Store › model › Embeddings
Qwen3.5-9B-DFlash-GGUF
by Anbeeld · source Hugging Face · updated 2026-08-29
apache-2.0730 MB~2 GB RAMsource aliveunlabeled
GGUF quantizations of z-lab DFlash draft model for Qwen 3.5 9B.
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:Anbeeld/Qwen3.5-9B-DFlash-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/Anbeeld/Qwen3.5-9B-DFlash-GGUF
- License: apache-2.0
- Requirements: about 2 GB of RAM, 730 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp.
- Tags:
transformersggufqwen3feature-extractionsafetensorsdflashspeculative-decodingspeculative-decoding-draftblock-diffusiondraft-modeldiffusion-language-modelefficiencyqwenqwen3.5sglangtext-generation
Numbers
- 4,712 downloads on Hugging Face
- 5 likes
- license apache-2.0
- 0.7 GB for qwen35-9b-dflash-Q4_K_M.gguf
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:59 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
qwen35-9b-dflash-Q2_K.gguf | Q2_K | 460 MB |
qwen35-9b-dflash-Q3_K_M.gguf | Q3_K_M | 595 MB |
qwen35-9b-dflash-Q4_K_M.gguf | Q4_K_M | 730 MB |
qwen35-9b-dflash-Q5_K_M.gguf | Q5_K_M | 871 MB |
qwen35-9b-dflash-Q6_K.gguf | Q6_K | 1021 MB |
qwen35-9b-dflash-Q8_0.gguf | Q8_0 | 1.3 GB |
qwen35-9b-dflash-bf16.gguf | BF16 | 2.4 GB |
From the source README
Qwen 3.5 9B DFlash GGUF
GGUF quantizations of z-lab DFlash draft model for Qwen 3.5 9B.
Use with BeeLlama.cpp, a llama.cpp fork with advanced quantization features.
Qwen3.5-9B-DFlash
Paper | Github | Blog
This DFlash draft model is a joint retrain from Z-Lab and Modal, trained with 40k sequence length and sliding-window attention for improved long-context performance. It is mirrored across the following Hugging Face repositories:
- `z-lab/Qwen3.5-9B-DFlash`
- `modal-labs/Qwen3.5-9B-DFlash`
This repository contains a DFlash draft model for `Qwen/Qwen3.5-9B`. It is not a standalone language model. It is intended to be paired with the target model in a speculative decoding server.
DFlash uses a lightweight block diffusion draft model to propose multiple tokens in parallel. The target model verifies those proposals, improving serving throughput while preserving the target model's output distribution.
Quick Start
Installation
SGLang
Card id model:hf:Anbeeld/Qwen3.5-9B-DFlash-GGUF · collected 2026-10-02 20:59 UTC · JSON