Store › model › Embeddings
Qwen3.6-35B-A3B-DFlash-GGUF
by Anbeeld · source Hugging Face · updated 2026-08-29
apache-2.0225 MB~1 GB RAMsource aliveunlabeled
GGUF quantizations of z-lab DFlash draft model for Qwen 3.6 35B A3B.
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:Anbeeld/Qwen3.6-35B-A3B-DFlash-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/Anbeeld/Qwen3.6-35B-A3B-DFlash-GGUF
- License: apache-2.0
- Requirements: about 1 GB of RAM, 225 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp.
- Tags:
transformersggufqwen3feature-extractionsafetensorsdflashspeculative-decodingspeculative-decoding-draftblock-diffusiondraft-modeldiffusion-language-modelefficiencyqwenqwen3.6sglangtext-generation
Numbers
- 4,444 downloads on Hugging Face
- 13 likes
- license apache-2.0
- 0.2 GB for qwen36-35b-a3b-dflash-Q4_K_M.gguf
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:59 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
qwen36-35b-a3b-dflash-Q2_K.gguf | Q2_K | 146 MB |
qwen36-35b-a3b-dflash-Q3_K_M.gguf | Q3_K_M | 187 MB |
qwen36-35b-a3b-dflash-Q4_K_M.gguf | Q4_K_M | 225 MB |
qwen36-35b-a3b-dflash-Q5_K_M.gguf | Q5_K_M | 267 MB |
qwen36-35b-a3b-dflash-Q6_K.gguf | Q6_K | 312 MB |
qwen36-35b-a3b-dflash-Q8_0.gguf | Q8_0 | 402 MB |
qwen36-35b-a3b-dflash-bf16.gguf | BF16 | 747 MB |
From the source README
Qwen 3.6 35B A3B DFlash GGUF
GGUF quantizations of z-lab DFlash draft model for Qwen 3.6 35B A3B.
Use with BeeLlama.cpp, a llama.cpp fork with advanced quantization features.
Qwen3.6-35B-A3B-DFlash
Paper | Github | Blog
This DFlash draft model is a joint retrain from Z-Lab and Modal, trained with 40k sequence length and sliding-window attention for improved long-context performance. It is mirrored across the following Hugging Face repositories:
- `z-lab/Qwen3.6-35B-A3B-DFlash`
- `modal-labs/Qwen3.6-35B-A3B-DFlash`
This repository contains a DFlash draft model for `Qwen/Qwen3.6-35B-A3B`. It is not a standalone language model. It is intended to be paired with the target model in a speculative decoding server.
DFlash uses a lightweight block diffusion draft model to propose multiple tokens in parallel. The target model verifies those proposals, improving serving throughput while preserving the target model's output distribution.
Quick Start
Installation
SGLang
Card id model:hf:Anbeeld/Qwen3.6-35B-A3B-DFlash-GGUF · collected 2026-10-02 20:59 UTC · JSON