{"v":1,"id":"model:hf:Anbeeld/Qwen3.5-9B-DFlash-GGUF","slug":"model-anbeeld-qwen3-5-9b-dflash-gguf","kind":"model","category":"embedding","title":"Qwen3.5-9B-DFlash-GGUF","summary":"GGUF quantizations of z-lab DFlash draft model for Qwen 3.5 9B.","source":{"provider":"hf","ref":"Anbeeld/Qwen3.5-9B-DFlash-GGUF","url":"https://huggingface.co/Anbeeld/Qwen3.5-9B-DFlash-GGUF","rev":"9bb5f9196eee5804fb7dd5c597726543989ee09f","fetchedAt":"2026-10-02T20:59:34.626Z","etag":"W/\"5009-rkRyCSTkhWUS+TJjYMk6PfVlfgI\""},"author":{"name":"Anbeeld","url":"https://huggingface.co/Anbeeld"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":4712,"downloadsWeek":327838,"likes":5,"takenAt":"2026-10-02T20:59:34.626Z"},"tags":["transformers","gguf","qwen3","feature-extraction","safetensors","dflash","speculative-decoding","speculative-decoding-draft","block-diffusion","draft-model","diffusion-language-model","efficiency","qwen","qwen3.5","sglang","text-generation","custom_code","text-generation-inference","text-embeddings-inference","endpoints_compatible","conversational"],"pipeline":"feature-extraction","links":{"github":"QwenLM/Qwen3","npm":"node-llama-cpp"},"updatedAt":"2026-08-29T18:06:49.000Z","collectedAt":"2026-10-02T20:59:34.626Z","review":{"numbers":["4,712 downloads on Hugging Face","5 likes","license apache-2.0","0.7 GB for qwen35-9b-dflash-Q4_K_M.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:59:34.626Z","http":200},"description":"# Qwen 3.5 9B DFlash GGUF\n\nGGUF quantizations of [**z-lab DFlash draft model**](https://huggingface.co/z-lab/Qwen3.5-9B-DFlash) for [**Qwen 3.5 9B**](https://huggingface.co/Qwen/Qwen3.5-9B).\n\nUse with [BeeLlama.cpp](https://github.com/Anbeeld/beellama.cpp), a llama.cpp fork with advanced quantization features.\n\n---\n\n# Qwen3.5-9B-DFlash\n\n[Paper](https://arxiv.org/abs/2602.06036) | [Github](https://github.com/z-lab/dflash) | [Blog](https://z-lab.ai/projects/dflash)\n\nThis DFlash draft model is a joint retrain from [Z-Lab](https://z-lab.ai) and [Modal](https://modal.com), trained with 40k sequence length and sliding-window attention for improved long-context performance. It is mirrored across the following Hugging Face repositories:\n\n- [`z-lab/Qwen3.5-9B-DFlash`](https://huggingface.co/z-lab/Qwen3.5-9B-DFlash)\n- [`modal-labs/Qwen3.5-9B-DFlash`](https://huggingface.co/modal-labs/Qwen3.5-9B-DFlash)\n\nThis repository contains a DFlash draft model for `Qwen/Qwen3.5-9B`. It is not a standalone language model. It is intended to be paired with the target model in a speculative decoding server.\n\nDFlash uses a lightweight block diffusion draft model to propose multiple tokens in parallel. The target model verifies those proposals, improving serving throughput while preserving the target model's output distribution.\n\n## Quick Start\n\n### Installation\n\n#### SGLang\n\nInstall a recent SGLang build with DFlash support:\n\n```bash\nuv pip install --upgrade \"sglang[all]\"\n```\n\nFor best performance on Blackwell GPUs, use an SGLang build that includes DFlash, FA4/TRT-LLM attention, and FlashInfer support.\n\n#### vLLM\n\nFor vLLM support, please refer to [vllm-project/vllm#40898](https://github.com/vllm-project/vllm/pull/40898). We will update the PR to make it merge-ready soon.\n\n### Launch Server\n\nThis model should be used with an inference server that supports DFlash speculative decoding. An example SGLang deployment is:\n\n```bash\nexp…\n\nSource: https://huggingface.co/Anbeeld/Qwen3.5-9B-DFlash-GGUF","install":{"kind":"model","hfId":"Anbeeld/Qwen3.5-9B-DFlash-GGUF","gated":false,"format":"gguf","files":[{"name":"qwen35-9b-dflash-Q2_K.gguf","size":481861312,"quant":"Q2_K","sha256":"650e6976095edaf1680f336675e85ef36c39ab69bfa97b1c985037dfdf5f9b8f"},{"name":"qwen35-9b-dflash-Q3_K_M.gguf","size":624139968,"quant":"Q3_K_M","sha256":"d6bb2bf80c128a2466866768ea4c988459aa7ea6ed1fde20c4b3957eae67b022"},{"name":"qwen35-9b-dflash-Q4_K_M.gguf","size":765959872,"quant":"Q4_K_M","sha256":"56bf9df07d6f8b140c818d387d4758530bee1ba18f10d54982dba602471035e7"},{"name":"qwen35-9b-dflash-Q5_K_M.gguf","size":913809088,"quant":"Q5_K_M","sha256":"51d010a41b6b5bd2f90370a221053da91a38530183baf1c2d0c8e60a1e4a336b"},{"name":"qwen35-9b-dflash-Q6_K.gguf","size":1070898880,"quant":"Q6_K","sha256":"42d8431c67da9c7db5c0a07811c1577a576c0dc9e817cb25685318eb165bbb81"},{"name":"qwen35-9b-dflash-Q8_0.gguf","size":1383767744,"quant":"Q8_0","sha256":"a745a9cbad9a9869954643aa24fb800068003c452c5b85c91958b8008aa44957"},{"name":"qwen35-9b-dflash-bf16.gguf","size":2594873024,"quant":"BF16","sha256":"28a485ab7ed4d7cfb6b788b7e481bd7d28785350a4b71bcb483e0e8b5da471bb"}],"totalBytes":7835309888,"suggestedFile":"qwen35-9b-dflash-Q4_K_M.gguf","requirements":{"ramGb":2,"diskBytes":765959872,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp"],"command":"lsh models install hf:Anbeeld/Qwen3.5-9B-DFlash-GGUF"}}