{"v":1,"id":"model:hf:empero-ai/Qwen3.8-4B-Distill-GGUF","slug":"model-empero-ai-qwen3-8-4b-distill-gguf","kind":"model","category":"llm","title":"Qwen3.8-4B-Distill-GGUF","summary":"GGUF quantizations of empero-ai/Qwen3.8-4B — a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-4B architecture — for llama.cpp, Ollama, LM Studio, Jan, KoboldCpp, and other stock GG…","source":{"provider":"hf","ref":"empero-ai/Qwen3.8-4B-Distill-GGUF","url":"https://huggingface.co/empero-ai/Qwen3.8-4B-Distill-GGUF","rev":"391fc7d103e3942a408def3e4f51c2f85d464417","fetchedAt":"2026-10-02T20:59:50.539Z","etag":"W/\"2a82-XH740FLFJYyNGY4Qisz3qBKJOIY\""},"author":{"name":"empero-ai","url":"https://huggingface.co/empero-ai"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":666186,"downloadsWeek":327838,"likes":161,"takenAt":"2026-10-02T20:59:50.539Z"},"tags":["gguf","llama.cpp","quantized","empero-ai","qwen3.5","qwen3.8","distillation","reasoning","gated-deltanet","text-generation","en","endpoints_compatible","conversational"],"pipeline":"text-generation","links":{"github":"QwenLM/Qwen3","npm":"node-llama-cpp"},"updatedAt":"2026-08-16T01:00:50.000Z","collectedAt":"2026-10-02T20:59:50.539Z","review":{"numbers":["666,186 downloads on Hugging Face","161 likes","license apache-2.0","2.6 GB for Qwen3.8-4B-Q4_K_M.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:59:50.539Z","http":200},"description":"# Qwen3.8-4B — GGUF\n\n**Developed by [Empero](https://empero.org)**\n\nGGUF quantizations of **[empero-ai/Qwen3.8-4B](https://huggingface.co/empero-ai/Qwen3.8-4B)** — a full-parameter distillation of **Qwen3.8 2.4T A95B** into the Qwen3.5-4B architecture — for [llama.cpp](https://github.com/ggml-org/llama.cpp), Ollama, LM Studio, Jan, KoboldCpp, and other stock GGUF runtimes.\n\nThis card is about choosing a file and running it. The capability writeup, full benchmark results, and best practices live on the **[main model card](https://huggingface.co/empero-ai/Qwen3.8-4B)**.\n\nHeadline results for the source model (CoT protocols, `lm-evaluation-harness`, identical settings base vs. student):\n\n| Task | Qwen3.5-4B (base) | **Qwen3.8-4B** | Δ |\n|---|---:|---:|---:|\n| mmlu (CoT, 57 subjects) | 0.354 | **0.553** | **+0.199** |\n| gsm8k_cot | 0.850 | 0.785 | −0.065 |\n\n> [!Note]\n> Qwen3.5-class models are hybrids: three Gated DeltaNet layers for every full-attention layer. A **recent llama.cpp build with Qwen3.5 / Gated DeltaNet support** is required — older builds will fail to load the architecture.\n\n## Files\n\n| File | Quant | Size | Notes |\n|---|---|---:|---|\n| `Qwen3.8-4B-Q4_K_M.gguf` | Q4_K_M | 2.783 GB | **Recommended.** Best quality/size balance for most users. |\n| `Qwen3.8-4B-Q5_K_M.gguf` | Q5_K_M | 3.161 GB | Higher quality at a modest size increase. |\n| `Qwen3.8-4B-Q6_K.gguf` | Q6_K | 3.563 GB | Near-lossless. |\n| `Qwen3.8-4B-Q8_0.gguf` | Q8_0 | 4.611 GB | Highest-quality quantization. |\n| `Qwen3.8-4B-BF16.gguf` | BF16 | 8.666 GB | Full precision reference. |\n\nSizes are exact decimal GB from the uploaded files (1 GB = 1,000,000,000 bytes).\n\n### What fits on a GPU?\n\nPractical weight-size-based guidance at modest context — the KV cache is the dominant cost at long context and may require offload regardless of weight quant:\n\n| Quant | Guidance |\n|---|---|\n| Q4_K_M / Q5_K_M | Comfortable on 4–6 GB cards; strong…\n\nSource: https://huggingface.co/empero-ai/Qwen3.8-4B-Distill-GGUF","install":{"kind":"model","hfId":"empero-ai/Qwen3.8-4B-Distill-GGUF","gated":false,"format":"gguf","files":[{"name":"Qwen3.8-4B-BF16.gguf","size":8665619744,"quant":"BF16","sha256":"448616595da523e57f694e1c8379aa5700bb1e3d9273eb2db4b9dab91bb85c6e"},{"name":"Qwen3.8-4B-Q4_K_M.gguf","size":2783446304,"quant":"Q4_K_M","sha256":"dec96e8cf2e11b613bb46513dec485377f9ca5a351e71712ee0e244f287c6790"},{"name":"Qwen3.8-4B-Q5_K_M.gguf","size":3161425184,"quant":"Q5_K_M","sha256":"735cd00b154f1a3f88899e7cc79e6a15b056b65d71049a75d88ec2948c4c0892"},{"name":"Qwen3.8-4B-Q6_K.gguf","size":3563027744,"quant":"Q6_K","sha256":"529393d9f7859122da727a8b662ea063127fb4320af8f58496a794b9bbf46e65"},{"name":"Qwen3.8-4B-Q8_0.gguf","size":4610579744,"quant":"Q8_0","sha256":"770b780d6754a4954d1caf395c9239eaeb394f15c7a7ea34039883377c93c9c3"}],"totalBytes":22784098720,"suggestedFile":"Qwen3.8-4B-Q4_K_M.gguf","requirements":{"ramGb":4,"diskBytes":2783446304,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp","ollama"],"command":"lsh models install hf:empero-ai/Qwen3.8-4B-Distill-GGUF"}}