{"v":1,"id":"model:hf:Anbeeld/Qwen3.5-4B-DFlash-GGUF","slug":"model-anbeeld-qwen3-5-4b-dflash-gguf","kind":"model","category":"embedding","title":"Qwen3.5-4B-DFlash-GGUF","summary":"GGUF quantizations of z-lab DFlash draft model for Qwen 3.5 4B.","source":{"provider":"hf","ref":"Anbeeld/Qwen3.5-4B-DFlash-GGUF","url":"https://huggingface.co/Anbeeld/Qwen3.5-4B-DFlash-GGUF","rev":"a79bafa7f383f8627f925250f547fe1ebbc43ff6","fetchedAt":"2026-10-02T21:00:17.626Z","etag":"W/\"5003-iuGibaU5dfINMLIy/iGjzy4Ubfk\""},"author":{"name":"Anbeeld","url":"https://huggingface.co/Anbeeld"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":2079,"downloadsWeek":327838,"likes":3,"takenAt":"2026-10-02T21:00:17.626Z"},"tags":["transformers","gguf","qwen3","feature-extraction","safetensors","dflash","speculative-decoding","speculative-decoding-draft","block-diffusion","draft-model","diffusion-language-model","efficiency","qwen","qwen3.5","sglang","text-generation","custom_code","text-generation-inference","text-embeddings-inference","endpoints_compatible","conversational"],"pipeline":"feature-extraction","links":{"github":"QwenLM/Qwen3","npm":"node-llama-cpp"},"updatedAt":"2026-08-29T18:07:42.000Z","collectedAt":"2026-10-02T21:00:17.626Z","review":{"numbers":["2,079 downloads on Hugging Face","3 likes","license apache-2.0","0.4 GB for qwen35-4b-dflash-Q4_K_M.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T21:00:17.626Z","http":200},"description":"# Qwen 3.5 4B DFlash GGUF\n\nGGUF quantizations of [**z-lab DFlash draft model**](https://huggingface.co/z-lab/Qwen3.5-4B-DFlash) for [**Qwen 3.5 4B**](https://huggingface.co/Qwen/Qwen3.5-4B).\n\nUse with [BeeLlama.cpp](https://github.com/Anbeeld/beellama.cpp), a llama.cpp fork with advanced quantization features.\n\n---\n\n# Qwen3.5-4B-DFlash\n\n[Paper](https://arxiv.org/abs/2602.06036) | [Github](https://github.com/z-lab/dflash) | [Blog](https://z-lab.ai/projects/dflash)\n\nThis DFlash draft model is a joint retrain from [Z-Lab](https://z-lab.ai) and [Modal](https://modal.com), trained with 40k sequence length and sliding-window attention for improved long-context performance. It is mirrored across the following Hugging Face repositories:\n\n- [`z-lab/Qwen3.5-4B-DFlash`](https://huggingface.co/z-lab/Qwen3.5-4B-DFlash)\n- [`modal-labs/Qwen3.5-4B-DFlash`](https://huggingface.co/modal-labs/Qwen3.5-4B-DFlash)\n\nThis repository contains a DFlash draft model for `Qwen/Qwen3.5-4B`. It is not a standalone language model. It is intended to be paired with the target model in a speculative decoding server.\n\nDFlash uses a lightweight block diffusion draft model to propose multiple tokens in parallel. The target model verifies those proposals, improving serving throughput while preserving the target model's output distribution.\n\n## Quick Start\n\n### Installation\n\n#### SGLang\n\nInstall a recent SGLang build with DFlash support:\n\n```bash\nuv pip install --upgrade \"sglang[all]\"\n```\n\nFor best performance on Blackwell GPUs, use an SGLang build that includes DFlash, FA4/TRT-LLM attention, and FlashInfer support.\n\n#### vLLM\n\nFor vLLM support, please refer to [vllm-project/vllm#40898](https://github.com/vllm-project/vllm/pull/40898). We will update the PR to make it merge-ready soon.\n\n### Launch Server\n\nThis model should be used with an inference server that supports DFlash speculative decoding. An example SGLang deployment is:\n\n```bash\nexp…\n\nSource: https://huggingface.co/Anbeeld/Qwen3.5-4B-DFlash-GGUF","install":{"kind":"model","hfId":"Anbeeld/Qwen3.5-4B-DFlash-GGUF","gated":false,"format":"gguf","files":[{"name":"qwen35-4b-dflash-Q2_K.gguf","size":243707584,"quant":"Q2_K","sha256":"d4da5ebcf35d13d1b05aabb381f75c583ab0fe1acdb24c57739ad28481c0fa5f"},{"name":"qwen35-4b-dflash-Q3_K_M.gguf","size":313585344,"quant":"Q3_K_M","sha256":"973e4eefe577e7afc8f778651b829c1d20326a5cd8831eec38282a993888dad9"},{"name":"qwen35-4b-dflash-Q4_K_M.gguf","size":381456064,"quant":"Q4_K_M","sha256":"83586a63eab58b4ef9dd1d72de667738fcedb289e97d98d01d8df66f23df610f"},{"name":"qwen35-4b-dflash-Q5_K_M.gguf","size":454201024,"quant":"Q5_K_M","sha256":"c10df5c5dd33456b478400b94f765aec2e5ff5e0920dee98e3b08b20928a20a8"},{"name":"qwen35-4b-dflash-Q6_K.gguf","size":531492544,"quant":"Q6_K","sha256":"a974f4471bde293f77587a4070e1f96fa8210af4bf79ad2edcf9303c7e249c06"},{"name":"qwen35-4b-dflash-Q8_0.gguf","size":685133504,"quant":"Q8_0","sha256":"9b6ffdd36bd9cb88a694461fe0aef68f6983ea1c3a79d63e0da552b326013a02"},{"name":"qwen35-4b-dflash-bf16.gguf","size":1279872704,"quant":"BF16","sha256":"18cefa06c73309385eff05eae315d9e91c0606e44f8994e61c08c63683894521"}],"totalBytes":3889448768,"suggestedFile":"qwen35-4b-dflash-Q4_K_M.gguf","requirements":{"ramGb":1,"diskBytes":381456064,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp"],"command":"lsh models install hf:Anbeeld/Qwen3.5-4B-DFlash-GGUF"}}