{"v":1,"id":"model:hf:Anbeeld/Qwen3.6-35B-A3B-DFlash-GGUF","slug":"model-anbeeld-qwen3-6-35b-a3b-dflash-gguf","kind":"model","category":"embedding","title":"Qwen3.6-35B-A3B-DFlash-GGUF","summary":"GGUF quantizations of z-lab DFlash draft model for Qwen 3.6 35B A3B.","source":{"provider":"hf","ref":"Anbeeld/Qwen3.6-35B-A3B-DFlash-GGUF","url":"https://huggingface.co/Anbeeld/Qwen3.6-35B-A3B-DFlash-GGUF","rev":"176b07a39a3de56f359a4a086064fe285165fed7","fetchedAt":"2026-10-02T20:59:36.632Z","etag":"W/\"507d-Gt9FXcYXjHBXMGuz2DUCKvb3v5g\""},"author":{"name":"Anbeeld","url":"https://huggingface.co/Anbeeld"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":4444,"downloadsWeek":327838,"likes":13,"takenAt":"2026-10-02T20:59:36.632Z"},"tags":["transformers","gguf","qwen3","feature-extraction","safetensors","dflash","speculative-decoding","speculative-decoding-draft","block-diffusion","draft-model","diffusion-language-model","efficiency","qwen","qwen3.6","sglang","text-generation","custom_code","text-generation-inference","text-embeddings-inference","endpoints_compatible","conversational"],"pipeline":"feature-extraction","links":{"github":"QwenLM/Qwen3","npm":"node-llama-cpp"},"updatedAt":"2026-08-29T18:08:08.000Z","collectedAt":"2026-10-02T20:59:36.632Z","review":{"numbers":["4,444 downloads on Hugging Face","13 likes","license apache-2.0","0.2 GB for qwen36-35b-a3b-dflash-Q4_K_M.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:59:36.632Z","http":200},"description":"# Qwen 3.6 35B A3B DFlash GGUF\n\nGGUF quantizations of [**z-lab DFlash draft model**](https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash) for [**Qwen 3.6 35B A3B**](https://huggingface.co/Qwen/Qwen3.6-35B-A3B).\n\nUse with [BeeLlama.cpp](https://github.com/Anbeeld/beellama.cpp), a llama.cpp fork with advanced quantization features.\n\n---\n\n# Qwen3.6-35B-A3B-DFlash\n\n[Paper](https://arxiv.org/abs/2602.06036) | [Github](https://github.com/z-lab/dflash) | [Blog](https://z-lab.ai/projects/dflash)\n\nThis DFlash draft model is a joint retrain from [Z-Lab](https://z-lab.ai) and [Modal](https://modal.com), trained with 40k sequence length and sliding-window attention for improved long-context performance. It is mirrored across the following Hugging Face repositories:\n\n- [`z-lab/Qwen3.6-35B-A3B-DFlash`](https://huggingface.co/z-lab/Qwen3.6-35B-A3B-DFlash)\n- [`modal-labs/Qwen3.6-35B-A3B-DFlash`](https://huggingface.co/modal-labs/Qwen3.6-35B-A3B-DFlash)\n\nThis repository contains a DFlash draft model for `Qwen/Qwen3.6-35B-A3B`. It is not a standalone language model. It is intended to be paired with the target model in a speculative decoding server.\n\nDFlash uses a lightweight block diffusion draft model to propose multiple tokens in parallel. The target model verifies those proposals, improving serving throughput while preserving the target model's output distribution.\n\n## Quick Start\n\n### Installation\n\n#### SGLang\n\nInstall a recent SGLang build with DFlash support:\n\n```bash\nuv pip install --upgrade \"sglang[all]\"\n```\n\nFor best performance on Blackwell GPUs, use an SGLang build that includes DFlash, FA4/TRT-LLM attention, and FlashInfer support.\n\n#### vLLM\n\nFor vLLM support, please refer to [vllm-project/vllm#40898](https://github.com/vllm-project/vllm/pull/40898). We will update the PR to make it merge-ready soon.\n\n### Launch Server\n\nThis model should be used with an inference server that supports DFlash speculative…\n\nSource: https://huggingface.co/Anbeeld/Qwen3.6-35B-A3B-DFlash-GGUF","install":{"kind":"model","hfId":"Anbeeld/Qwen3.6-35B-A3B-DFlash-GGUF","gated":false,"format":"gguf","files":[{"name":"qwen36-35b-a3b-dflash-Q2_K.gguf","size":153411296,"quant":"Q2_K","sha256":"21221ae8d66151549512bafc84a76193566991497e15ccbd0f533e817ad92f5a"},{"name":"qwen36-35b-a3b-dflash-Q3_K_M.gguf","size":195780320,"quant":"Q3_K_M","sha256":"2a16c83211a6acda218a8dc6b0f389342c1232d7c818544046c0030e01b9edf4"},{"name":"qwen36-35b-a3b-dflash-Q4_K_M.gguf","size":235691744,"quant":"Q4_K_M","sha256":"b7dd951da39555463537c1838e914d9ff4294edae463c5bb0b16f45ffc5d0a8c"},{"name":"qwen36-35b-a3b-dflash-Q5_K_M.gguf","size":280256224,"quant":"Q5_K_M","sha256":"91f63edf70fe7d9fa1dbd51f21791d4a3e99fd87c00091a70ed93e7415952e5b"},{"name":"qwen36-35b-a3b-dflash-Q6_K.gguf","size":327605984,"quant":"Q6_K","sha256":"a6d76c32b897715cabadb0ee75760d8c3fea23f38652a53ee0ce0b0fb2cc5c7d"},{"name":"qwen36-35b-a3b-dflash-Q8_0.gguf","size":421060320,"quant":"Q8_0","sha256":"b9749d1793b8c90706168ec6b7cfedee74d3c76e67c4c61cf477e08ea478761c"},{"name":"qwen36-35b-a3b-dflash-bf16.gguf","size":782819040,"quant":"BF16","sha256":"8d247ef12df757af77e8fcca7eb1771835b66de1c5d77869b6e12174eba06cde"}],"totalBytes":2396624928,"suggestedFile":"qwen36-35b-a3b-dflash-Q4_K_M.gguf","requirements":{"ramGb":1,"diskBytes":235691744,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp"],"command":"lsh models install hf:Anbeeld/Qwen3.6-35B-A3B-DFlash-GGUF"}}