{"v":1,"id":"model:hf:z-lab/Qwen3.8-27B-DFlash2-GGUF","slug":"model-z-lab-qwen3-8-27b-dflash2-gguf","kind":"model","category":"llm","title":"Qwen3.8-27B-DFlash2-GGUF","summary":"This repository contains GGUF conversions of incoai/Qwen3.8-27B-DFlash2, the DFlash 2 draft model for Qwen/Qwen3.8-27B. It is not a standalone language model: it runs inside a speculative decoding se…","source":{"provider":"hf","ref":"z-lab/Qwen3.8-27B-DFlash2-GGUF","url":"https://huggingface.co/z-lab/Qwen3.8-27B-DFlash2-GGUF","rev":"2d9571f8ce46e151f61c6499c99dee6079e1d610","fetchedAt":"2026-10-02T20:53:05.513Z","etag":"W/\"2daf-9HIx26uuj/NnKvpMQ78xEdvtHHU\""},"author":{"name":"z-lab","url":"https://huggingface.co/z-lab"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":520168,"downloadsWeek":327838,"likes":146,"stars":27665,"openIssues":68,"pushedAt":"2026-01-09T03:05:47Z","takenAt":"2026-10-02T20:53:05.513Z"},"tags":["llama.cpp","gguf","dflash2","speculative-decoding","draft-model","text-generation","conversational"],"pipeline":"text-generation","links":{"github":"QwenLM/Qwen3","npm":"node-llama-cpp"},"updatedAt":"2026-08-24T23:07:28.000Z","collectedAt":"2026-10-02T20:53:05.513Z","review":{"numbers":["520,168 downloads on Hugging Face","146 likes","license apache-2.0","1.1 GB for Qwen3.8-27B-DFlash2-Q4_K_M.gguf","27,665 stars on QwenLM/Qwen3","68 open issues and PRs","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:53:05.513Z","http":200},"description":"# Qwen3.8-27B-DFlash2-GGUF\n\n[Blog](https://inco.ai/blog/dflash2/) | [GitHub](https://github.com/z-lab/dflash)\n\nThis repository contains GGUF conversions of\n[`incoai/Qwen3.8-27B-DFlash2`](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2),\nthe DFlash 2 draft model for\n[`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B).\nIt is not a standalone language model: it runs inside a speculative\ndecoding server and drafts tokens for the target model to verify. This\nrepository is a mirror of\n[`incoai/Qwen3.8-27B-DFlash2-GGUF`](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2-GGUF).\n\nDFlash 2 is a block-diffusion drafter for speculative decoding. It predicts\na whole block of tokens in a single pass and keeps the top candidates at\nevery position. A lightweight selector then traces one coherent path through\nthem. Two-tap dynamic convolutions in the backbone keep the draft from\ndecaying toward the end of the block. Decoding is lossless: greedy output\nmatches the target model exactly, and sampling preserves its distribution.\n\n  \n\n| File | Size |\n| :--- | ---: |\n| `Qwen3.8-27B-DFlash2-Q4_K_M.gguf` | 1.1 GB |\n| `Qwen3.8-27B-DFlash2-Q8_0.gguf` | 2.0 GB |\n| `Qwen3.8-27B-DFlash2-BF16.gguf` | 3.8 GB |\n\n## Quick Start\n\nBuild [llama.cpp](https://github.com/ggml-org/llama.cpp) with DFlash 2\nsupport ([PR #27342](https://github.com/ggml-org/llama.cpp/pull/27342)):\n\n```bash\ngit clone https://github.com/ggml-org/llama.cpp.git\ncd llama.cpp\ngit fetch origin pull/27342/head:pr-27342\ngit switch pr-27342\n\n# NVIDIA CUDA\ncmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON\ncmake --build build -j\n\n# Apple Silicon\ncmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON\ncmake --build build -j\n```\n\nThen serve:\n\n```bash\n./build/bin/llama-server \\\n  -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M \\\n  -hfd incoai/Qwen3.8-27B-DFlash2-GGUF:Q4_K_M \\\n  --spec-type draft-dflash \\\n  --spec-draft-n-max 7\n```\n\nSee the [blog post](https://inco.ai/b…\n\nSource: https://huggingface.co/z-lab/Qwen3.8-27B-DFlash2-GGUF","install":{"kind":"model","hfId":"z-lab/Qwen3.8-27B-DFlash2-GGUF","gated":false,"format":"gguf","files":[{"name":"Qwen3.8-27B-DFlash2-BF16.gguf","size":3860293216,"quant":"BF16","sha256":"26d47ca20ab07688327a63d912acad222d924eaaa92a980cc488de3c67e736bc"},{"name":"Qwen3.8-27B-DFlash2-Q4_K_M.gguf","size":1143006816,"quant":"Q4_K_M","sha256":"1a25c56858e1ebe93f2718ac1d49d1151f9323325c1bbfd6209370f4db131ebd"},{"name":"Qwen3.8-27B-DFlash2-Q8_0.gguf","size":2056414816,"quant":"Q8_0","sha256":"c18e800daedc59ca68fd13b6a856d795746af6d399a9279ac6a277d1d422f87e"}],"totalBytes":7059714848,"suggestedFile":"Qwen3.8-27B-DFlash2-Q4_K_M.gguf","requirements":{"ramGb":2,"diskBytes":1143006816,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp","ollama"],"command":"lsh models install hf:z-lab/Qwen3.8-27B-DFlash2-GGUF"}}