{"v":1,"id":"model:hf:empero-ai/Qwen3.8-9B-Distill-GGUF","slug":"model-empero-ai-qwen3-8-9b-distill-gguf","kind":"model","category":"llm","title":"Qwen3.8-9B-Distill-GGUF","summary":"GGUF quantizations of empero-ai/Qwen3.8-9B — a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-9B architecture — for llama.cpp, Ollama, LM Studio, Jan, KoboldCpp, and other stock GG…","source":{"provider":"hf","ref":"empero-ai/Qwen3.8-9B-Distill-GGUF","url":"https://huggingface.co/empero-ai/Qwen3.8-9B-Distill-GGUF","rev":"760121cd70bb4c36b2b5ec58eb765e0df5987efe","fetchedAt":"2026-10-02T20:59:44.409Z","etag":"W/\"2a85-NMsVJE6BdLoGjFcUMZp57A923S4\""},"author":{"name":"empero-ai","url":"https://huggingface.co/empero-ai"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":675501,"downloadsWeek":327838,"likes":286,"takenAt":"2026-10-02T20:59:44.409Z"},"tags":["gguf","llama.cpp","quantized","empero-ai","qwen3.5","qwen3.8","distillation","reasoning","gated-deltanet","text-generation","en","endpoints_compatible","conversational"],"pipeline":"text-generation","links":{"github":"QwenLM/Qwen3","npm":"node-llama-cpp"},"updatedAt":"2026-08-16T01:00:49.000Z","collectedAt":"2026-10-02T20:59:44.409Z","review":{"numbers":["675,501 downloads on Hugging Face","286 likes","license apache-2.0","5.4 GB for Qwen3.8-9B-Q4_K_M.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:59:44.409Z","http":200},"description":"# Qwen3.8-9B — GGUF\n\n**Developed by [Empero](https://empero.org)**\n\nGGUF quantizations of **[empero-ai/Qwen3.8-9B](https://huggingface.co/empero-ai/Qwen3.8-9B)** — a full-parameter distillation of **Qwen3.8 2.4T A95B** into the Qwen3.5-9B architecture — for [llama.cpp](https://github.com/ggml-org/llama.cpp), Ollama, LM Studio, Jan, KoboldCpp, and other stock GGUF runtimes.\n\nThis card is about choosing a file and running it. The capability writeup, full benchmark results, and best practices live on the **[main model card](https://huggingface.co/empero-ai/Qwen3.8-9B)**.\n\nHeadline results for the source model (CoT protocols, `lm-evaluation-harness`, identical settings base vs. student):\n\n| Task | Qwen3.5-9B (base) | **Qwen3.8-9B** | Δ |\n|---|---:|---:|---:|\n| mmlu (CoT, 57 subjects) | 0.546 | **0.751** | **+0.205** |\n| gsm8k_cot | 0.885 | 0.870 | −0.015 |\n\n> [!Note]\n> Qwen3.5-class models are hybrids: three Gated DeltaNet layers for every full-attention layer. A **recent llama.cpp build with Qwen3.5 / Gated DeltaNet support** is required — older builds will fail to load the architecture.\n\n## Files\n\n| File | Quant | Size | Notes |\n|---|---|---:|---|\n| `Qwen3.8-9B-Q4_K_M.gguf` | Q4_K_M | 5.780 GB | **Recommended.** Best quality/size balance for most users. |\n| `Qwen3.8-9B-Q5_K_M.gguf` | Q5_K_M | 6.643 GB | Higher quality, still fits an 8 GB card at short context. |\n| `Qwen3.8-9B-Q6_K.gguf` | Q6_K | 7.559 GB | Near-lossless. |\n| `Qwen3.8-9B-Q8_0.gguf` | Q8_0 | 9.786 GB | Highest-quality quantization. |\n| `Qwen3.8-9B-BF16.gguf` | BF16 | 18.407 GB | Full precision reference. |\n\nSizes are exact decimal GB from the uploaded files (1 GB = 1,000,000,000 bytes).\n\n### What fits on a GPU?\n\nPractical weight-size-based guidance at modest context — the KV cache is the dominant cost at long context and may require offload regardless of weight quant:\n\n| Quant | Guidance |\n|---|---|\n| Q4_K_M / Q5_K_M | Comfortable on 8–1…\n\nSource: https://huggingface.co/empero-ai/Qwen3.8-9B-Distill-GGUF","install":{"kind":"model","hfId":"empero-ai/Qwen3.8-9B-Distill-GGUF","gated":false,"format":"gguf","files":[{"name":"Qwen3.8-9B-BF16.gguf","size":18407320896,"quant":"BF16","sha256":"f2aee1994144502b6cf6c7bf4db30b39f2c2b67ce0069f831146bf1fe1e0517c"},{"name":"Qwen3.8-9B-Q4_K_M.gguf","size":5780090176,"quant":"Q4_K_M","sha256":"df13d66021cef676f82be74053220fd75af6bf2a6a7fb77f5222ab9e50744a7a"},{"name":"Qwen3.8-9B-Q5_K_M.gguf","size":6642543936,"quant":"Q5_K_M","sha256":"c6667345d4e45d8cddca3c8e997a483a4f9ae04ee9402273c24e959dee7173dc"},{"name":"Qwen3.8-9B-Q6_K.gguf","size":7558901056,"quant":"Q6_K","sha256":"0f1271373f899912bfe4ea76299af7dd83722d98ea421b0827501c3a2c6da22b"},{"name":"Qwen3.8-9B-Q8_0.gguf","size":9786060096,"quant":"Q8_0","sha256":"79ca5d342a07922f2bbf38c8d892a79a3c8620c65feaf4b1c66b7830ae724db8"}],"totalBytes":48174916160,"suggestedFile":"Qwen3.8-9B-Q4_K_M.gguf","requirements":{"ramGb":7,"diskBytes":5780090176,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp","ollama"],"command":"lsh models install hf:empero-ai/Qwen3.8-9B-Distill-GGUF"}}