Store › model › LLM
Qwen3.8-27B-DFlash2-GGUF
by z-lab · source Hugging Face · updated 2026-08-24
apache-2.01.1 GB~2 GB RAMsource aliveunlabeled
This repository contains GGUF conversions of incoai/Qwen3.8-27B-DFlash2, the DFlash 2 draft model for Qwen/Qwen3.8-27B. It is not a standalone language model: it runs inside a speculative decoding se…
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:z-lab/Qwen3.8-27B-DFlash2-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/z-lab/Qwen3.8-27B-DFlash2-GGUF
- License: apache-2.0
- Requirements: about 2 GB of RAM, 1.1 GB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp, ollama.
- Tags:
llama.cppggufdflash2speculative-decodingdraft-modeltext-generationconversational
Numbers
- 520,168 downloads on Hugging Face
- 146 likes
- license apache-2.0
- 1.1 GB for Qwen3.8-27B-DFlash2-Q4_K_M.gguf
- 27,665 stars on QwenLM/Qwen3
- 68 open issues and PRs
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:53 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
Qwen3.8-27B-DFlash2-BF16.gguf | BF16 | 3.6 GB |
Qwen3.8-27B-DFlash2-Q4_K_M.gguf | Q4_K_M | 1.1 GB |
Qwen3.8-27B-DFlash2-Q8_0.gguf | Q8_0 | 1.9 GB |
From the source README
Qwen3.8-27B-DFlash2-GGUF
Blog | GitHub
This repository contains GGUF conversions of
`incoai/Qwen3.8-27B-DFlash2`,
the DFlash 2 draft model for
`Qwen/Qwen3.8-27B`.
It is not a standalone language model: it runs inside a speculative
decoding server and drafts tokens for the target model to verify. This
repository is a mirror of
`incoai/Qwen3.8-27B-DFlash2-GGUF`.
DFlash 2 is a block-diffusion drafter for speculative decoding. It predicts
a whole block of tokens in a single pass and keeps the top candidates at
every position. A lightweight selector then traces one coherent path through
them. Two-tap dynamic convolutions in the backbone keep the draft from
decaying toward the end of the block. Decoding is lossless: greedy output
matches the target model exactly, and sampling preserves its distribution.
| File | Size |
| :--- | ---: |
| `Qwen3.8-27B-DFlash2-Q4_K_M.gguf` | 1.1 GB |
| `Qwen3.8-27B-DFlash2-Q8_0.gguf` | 2.0 GB |
| `Qwen3.8-27B-DFlash2-BF16.gguf` | 3.8 GB |
Quick Start
Build llama.cpp with DFlash 2
support (PR #27342):
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
git fetch origin pull/27342/head:pr-27342
git switch pr-27342
# NVIDIA CUDA
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON
cmake --build build -j
# Apple Silicon
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON
cmake --build build -j
```
Then serve:
./build/bin/llama-server \
-hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M \
-hfd incoai/Qwen3.8-27B-DFlash2-GGUF:Q4_K_M \
--spec-type draft-dflash \
--spec-draft-n-max 7
Card id model:hf:z-lab/Qwen3.8-27B-DFlash2-GGUF · collected 2026-10-02 20:53 UTC · JSON