LogiShell store Open app

Store › model › LLM

Qwen3.8-27B-DFlash2-GGUF

by z-lab · source Hugging Face · updated 2026-08-24

apache-2.01.1 GB~2 GB RAMsource aliveunlabeled

This repository contains GGUF conversions of incoai/Qwen3.8-27B-DFlash2, the DFlash 2 draft model for Qwen/Qwen3.8-27B. It is not a standalone language model: it runs inside a speculative decoding se…

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:z-lab/Qwen3.8-27B-DFlash2-GGUF and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-02 20:53 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
Qwen3.8-27B-DFlash2-BF16.ggufBF163.6 GB
Qwen3.8-27B-DFlash2-Q4_K_M.ggufQ4_K_M1.1 GB
Qwen3.8-27B-DFlash2-Q8_0.ggufQ8_01.9 GB

From the source README

Qwen3.8-27B-DFlash2-GGUF

Blog | GitHub

This repository contains GGUF conversions of
`incoai/Qwen3.8-27B-DFlash2`,
the DFlash 2 draft model for
`Qwen/Qwen3.8-27B`.
It is not a standalone language model: it runs inside a speculative
decoding server and drafts tokens for the target model to verify. This
repository is a mirror of
`incoai/Qwen3.8-27B-DFlash2-GGUF`.

DFlash 2 is a block-diffusion drafter for speculative decoding. It predicts
a whole block of tokens in a single pass and keeps the top candidates at
every position. A lightweight selector then traces one coherent path through
them. Two-tap dynamic convolutions in the backbone keep the draft from
decaying toward the end of the block. Decoding is lossless: greedy output
matches the target model exactly, and sampling preserves its distribution.

| File | Size |
| :--- | ---: |
| `Qwen3.8-27B-DFlash2-Q4_K_M.gguf` | 1.1 GB |
| `Qwen3.8-27B-DFlash2-Q8_0.gguf` | 2.0 GB |
| `Qwen3.8-27B-DFlash2-BF16.gguf` | 3.8 GB |

Quick Start

Build llama.cpp with DFlash 2
support (PR #27342):

git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
git fetch origin pull/27342/head:pr-27342
git switch pr-27342

# NVIDIA CUDA
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON
cmake --build build -j

# Apple Silicon
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON
cmake --build build -j
```

Then serve:

./build/bin/llama-server \
  -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M \
  -hfd incoai/Qwen3.8-27B-DFlash2-GGUF:Q4_K_M \
  --spec-type draft-dflash \
  --spec-draft-n-max 7

Card id model:hf:z-lab/Qwen3.8-27B-DFlash2-GGUF · collected 2026-10-02 20:53 UTC · JSON