LogiShell store Open app

Store › model › Embeddings

Qwen3.5-4B-DFlash-GGUF

by Anbeeld · source Hugging Face · updated 2026-08-29

apache-2.0364 MB~1 GB RAMsource aliveunlabeled

GGUF quantizations of z-lab DFlash draft model for Qwen 3.5 4B.

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:Anbeeld/Qwen3.5-4B-DFlash-GGUF and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-02 21:00 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
qwen35-4b-dflash-Q2_K.ggufQ2_K232 MB
qwen35-4b-dflash-Q3_K_M.ggufQ3_K_M299 MB
qwen35-4b-dflash-Q4_K_M.ggufQ4_K_M364 MB
qwen35-4b-dflash-Q5_K_M.ggufQ5_K_M433 MB
qwen35-4b-dflash-Q6_K.ggufQ6_K507 MB
qwen35-4b-dflash-Q8_0.ggufQ8_0653 MB
qwen35-4b-dflash-bf16.ggufBF161.2 GB

From the source README

Qwen 3.5 4B DFlash GGUF

GGUF quantizations of z-lab DFlash draft model for Qwen 3.5 4B.

Use with BeeLlama.cpp, a llama.cpp fork with advanced quantization features.

Qwen3.5-4B-DFlash

Paper | Github | Blog

This DFlash draft model is a joint retrain from Z-Lab and Modal, trained with 40k sequence length and sliding-window attention for improved long-context performance. It is mirrored across the following Hugging Face repositories:

  • `z-lab/Qwen3.5-4B-DFlash`
  • `modal-labs/Qwen3.5-4B-DFlash`

This repository contains a DFlash draft model for `Qwen/Qwen3.5-4B`. It is not a standalone language model. It is intended to be paired with the target model in a speculative decoding server.

DFlash uses a lightweight block diffusion draft model to propose multiple tokens in parallel. The target model verifies those proposals, improving serving throughput while preserving the target model's output distribution.

Quick Start

Installation

SGLang

Card id model:hf:Anbeeld/Qwen3.5-4B-DFlash-GGUF · collected 2026-10-02 21:00 UTC · JSON