{"v":1,"id":"model:hf:empero-ai/Qwen3.8-2B-Distill-GGUF","slug":"model-empero-ai-qwen3-8-2b-distill-gguf","kind":"model","category":"llm","title":"Qwen3.8-2B-Distill-GGUF","summary":"GGUF quantizations of empero-ai/Qwen3.8-2B — a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-2B architecture, the smallest member of the family — for llama.cpp, Ollama, LM Studio,…","source":{"provider":"hf","ref":"empero-ai/Qwen3.8-2B-Distill-GGUF","url":"https://huggingface.co/empero-ai/Qwen3.8-2B-Distill-GGUF","rev":"f4f73582d0b149595450c719b9a7521a03894f9c","fetchedAt":"2026-10-02T20:59:59.531Z","etag":"W/\"2a8f-vChZix2doHEjvG5Af6VllvLTXeA\""},"author":{"name":"empero-ai","url":"https://huggingface.co/empero-ai"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":621812,"downloadsWeek":327838,"likes":158,"takenAt":"2026-10-02T20:59:59.531Z"},"tags":["gguf","llama.cpp","quantized","empero-ai","qwen3.5","qwen3.8","distillation","reasoning","gated-deltanet","edge","text-generation","en","endpoints_compatible","conversational"],"pipeline":"text-generation","links":{"github":"QwenLM/Qwen3","npm":"node-llama-cpp"},"updatedAt":"2026-08-16T01:00:50.000Z","collectedAt":"2026-10-02T20:59:59.531Z","review":{"numbers":["621,812 downloads on Hugging Face","158 likes","license apache-2.0","1.2 GB for Qwen3.8-2B-Q4_K_M.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:59:59.531Z","http":200},"description":"# Qwen3.8-2B — GGUF\n\n**Developed by [Empero](https://empero.org)**\n\nGGUF quantizations of **[empero-ai/Qwen3.8-2B](https://huggingface.co/empero-ai/Qwen3.8-2B)** — a full-parameter distillation of **Qwen3.8 2.4T A95B** into the Qwen3.5-2B architecture, the smallest member of the family — for [llama.cpp](https://github.com/ggml-org/llama.cpp), Ollama, LM Studio, Jan, KoboldCpp, and other stock GGUF runtimes.\n\nThis card is about choosing a file and running it. The capability writeup, full benchmark results, and best practices live on the **[main model card](https://huggingface.co/empero-ai/Qwen3.8-2B)**.\n\nHeadline results for the source model (CoT protocols, `lm-evaluation-harness`, identical settings base vs. student):\n\n| Task | Qwen3.5-2B (base) | **Qwen3.8-2B** | Δ |\n|---|---:|---:|---:|\n| mmlu (CoT, 57 subjects) | 0.283 | **0.548** | **+0.265** |\n| gsm8k_cot | 0.330 | **0.640** | **+0.310** |\n\n> [!Note]\n> Qwen3.5-class models are hybrids: three Gated DeltaNet layers for every full-attention layer. A **recent llama.cpp build with Qwen3.5 / Gated DeltaNet support** is required — older builds will fail to load the architecture.\n\n## Files\n\n| File | Quant | Size | Notes |\n|---|---|---:|---|\n| `Qwen3.8-2B-Q4_K_M.gguf` | Q4_K_M | 1.312 GB | **Recommended.** Best quality/size balance; runs on phones and SBCs. |\n| `Qwen3.8-2B-Q5_K_M.gguf` | Q5_K_M | 1.455 GB | Higher quality at a modest size increase. |\n| `Qwen3.8-2B-Q6_K.gguf` | Q6_K | 1.606 GB | Near-lossless. |\n| `Qwen3.8-2B-Q8_0.gguf` | Q8_0 | 2.077 GB | Highest-quality quantization. |\n| `Qwen3.8-2B-BF16.gguf` | BF16 | 3.897 GB | Full precision reference. |\n\nSizes are exact decimal GB from the uploaded files (1 GB = 1,000,000,000 bytes).\n\n### Where it runs\n\nPractical weight-size-based guidance at modest context — the KV cache is the dominant cost at long context:\n\n| Quant | Guidance |\n|---|---|\n| Q4_K_M / Q5_K_M | Phones, single-board computers, any mod…\n\nSource: https://huggingface.co/empero-ai/Qwen3.8-2B-Distill-GGUF","install":{"kind":"model","hfId":"empero-ai/Qwen3.8-2B-Distill-GGUF","gated":false,"format":"gguf","files":[{"name":"Qwen3.8-2B-BF16.gguf","size":3897387392,"quant":"BF16","sha256":"44763f3d83f0a1a3ee63334b60916705dc565d796cb0f2b8c320414c57f4ac48"},{"name":"Qwen3.8-2B-Q4_K_M.gguf","size":1312164224,"quant":"Q4_K_M","sha256":"4aa0fb13c431514262f259d420ecc95a8714df58ac2a2384514e20b93983f0ff"},{"name":"Qwen3.8-2B-Q5_K_M.gguf","size":1454786944,"quant":"Q5_K_M","sha256":"609bbe7b681303356db22220abb17eed531e10c7f87d83e828788ea8693c8e5d"},{"name":"Qwen3.8-2B-Q6_K.gguf","size":1606323584,"quant":"Q6_K","sha256":"0c9fc69b74d8be52ee28c388518f31787b65be35d3db882a66a0d548eba8d4df"},{"name":"Qwen3.8-2B-Q8_0.gguf","size":2076674432,"quant":"Q8_0","sha256":"866773b0d68f09a1db9733555e92daff85b617f9a2e601773dff494c5ca2bbf2"}],"totalBytes":10347336576,"suggestedFile":"Qwen3.8-2B-Q4_K_M.gguf","requirements":{"ramGb":2,"diskBytes":1312164224,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp","ollama"],"command":"lsh models install hf:empero-ai/Qwen3.8-2B-Distill-GGUF"}}