LogiShell store Open app

Store › model › Embeddings

Qwen3.6-35B-A3B-DFlash-GGUF

by Anbeeld · source Hugging Face · updated 2026-08-29

apache-2.0225 MB~1 GB RAMsource aliveunlabeled

GGUF quantizations of z-lab DFlash draft model for Qwen 3.6 35B A3B.

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:Anbeeld/Qwen3.6-35B-A3B-DFlash-GGUF and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-02 20:59 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
qwen36-35b-a3b-dflash-Q2_K.ggufQ2_K146 MB
qwen36-35b-a3b-dflash-Q3_K_M.ggufQ3_K_M187 MB
qwen36-35b-a3b-dflash-Q4_K_M.ggufQ4_K_M225 MB
qwen36-35b-a3b-dflash-Q5_K_M.ggufQ5_K_M267 MB
qwen36-35b-a3b-dflash-Q6_K.ggufQ6_K312 MB
qwen36-35b-a3b-dflash-Q8_0.ggufQ8_0402 MB
qwen36-35b-a3b-dflash-bf16.ggufBF16747 MB

From the source README

Qwen 3.6 35B A3B DFlash GGUF

GGUF quantizations of z-lab DFlash draft model for Qwen 3.6 35B A3B.

Use with BeeLlama.cpp, a llama.cpp fork with advanced quantization features.

Qwen3.6-35B-A3B-DFlash

Paper | Github | Blog

This DFlash draft model is a joint retrain from Z-Lab and Modal, trained with 40k sequence length and sliding-window attention for improved long-context performance. It is mirrored across the following Hugging Face repositories:

  • `z-lab/Qwen3.6-35B-A3B-DFlash`
  • `modal-labs/Qwen3.6-35B-A3B-DFlash`

This repository contains a DFlash draft model for `Qwen/Qwen3.6-35B-A3B`. It is not a standalone language model. It is intended to be paired with the target model in a speculative decoding server.

DFlash uses a lightweight block diffusion draft model to propose multiple tokens in parallel. The target model verifies those proposals, improving serving throughput while preserving the target model's output distribution.

Quick Start

Installation

SGLang

Card id model:hf:Anbeeld/Qwen3.6-35B-A3B-DFlash-GGUF · collected 2026-10-02 20:59 UTC · JSON