{"v":1,"id":"model:hf:empero-ai/Qwen3.8-35B-A3B-Distill-GGUF","slug":"model-empero-ai-qwen3-8-35b-a3b-distill-gguf","kind":"model","category":"llm","title":"Qwen3.8-35B-A3B-Distill-GGUF","summary":"GGUF quantizations of empero-ai/Qwen3.8-35B-A3B-Distill — a distillation of the Qwen3.8 frontier models into the Qwen3.6-35B-A3B Mixture-of-Experts architecture — for llama.cpp, Ollama, LM Studio, Ja…","source":{"provider":"hf","ref":"empero-ai/Qwen3.8-35B-A3B-Distill-GGUF","url":"https://huggingface.co/empero-ai/Qwen3.8-35B-A3B-Distill-GGUF","rev":"b1f9d1dcc3de8aa867669b0ab919384aeeb9b8d5","fetchedAt":"2026-10-02T21:00:29.535Z","etag":"W/\"3081-eb8iCe130znYnebT1bzLvUTmsFE\""},"author":{"name":"empero-ai","url":"https://huggingface.co/empero-ai"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":479410,"downloadsWeek":327838,"likes":169,"takenAt":"2026-10-02T21:00:29.535Z"},"tags":["gguf","llama.cpp","quantized","empero-ai","qwen3.6","qwen3.8","distillation","reasoning","moe","gated-deltanet","text-generation","en","endpoints_compatible","conversational"],"pipeline":"text-generation","links":{"github":"QwenLM/Qwen3","npm":"node-llama-cpp"},"updatedAt":"2026-09-16T21:28:58.000Z","collectedAt":"2026-10-02T21:00:29.535Z","review":{"numbers":["479,410 downloads on Hugging Face","169 likes","license apache-2.0","20 GB for Qwen3.8-35B-A3B-Q4_K_M.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T21:00:29.535Z","http":200},"description":"# Qwen3.8-35B-A3B — GGUF\n\n**Developed by [Empero](https://empero.org)**\n\nGGUF quantizations of **[empero-ai/Qwen3.8-35B-A3B-Distill](https://huggingface.co/empero-ai/Qwen3.8-35B-A3B-Distill)** — a distillation of the Qwen3.8 frontier models into the Qwen3.6-35B-A3B Mixture-of-Experts architecture — for [llama.cpp](https://github.com/ggml-org/llama.cpp), Ollama, LM Studio, Jan, KoboldCpp, and other stock GGUF runtimes.\n\nThis card is about choosing a file and running it. The capability writeup, benchmark results, and best practices live on the **[main model card](https://huggingface.co/empero-ai/Qwen3.8-35B-A3B-Distill)**.\n\n35B total parameters with ~3B active per token — the MoE sparsity means it runs considerably faster than a dense 35B at the same quant, but the **whole weight file still has to fit in RAM or VRAM**.\n\n> [!Note]\n> Qwen3.6-class models are hybrids: 30 Gated DeltaNet layers and 10 full-attention layers, with 256 experts routed 8-per-token. A **recent llama.cpp build with Qwen3.6 / Gated DeltaNet MoE support** is required — older builds will fail to load the architecture.\n\n## Files\n\n| File | Quant | Size | Notes |\n|---|---|---:|---|\n| `Qwen3.8-35B-A3B-IQ2_M.gguf` | IQ2_M | 12.558 GB | Smallest usable. Fits a 16 GB card. |\n| `Qwen3.8-35B-A3B-Q2_K.gguf` | Q2_K | 13.839 GB | 2-bit K-quant; widest runtime support at this size. |\n| `Qwen3.8-35B-A3B-IQ3_M.gguf` | IQ3_M | 16.340 GB | Strong quality per byte at 3-bit. |\n| `Qwen3.8-35B-A3B-Q3_K_M.gguf` | Q3_K_M | 17.664 GB | Conventional 3-bit K-quant. |\n| `Qwen3.8-35B-A3B-IQ4_XS.gguf` | IQ4_XS | 19.628 GB | Near Q4_K_M quality, ~2 GB smaller. |\n| `Qwen3.8-35B-A3B-Q4_K_M.gguf` | Q4_K_M | 21.713 GB | **Recommended.** Best quality/size balance for most users. |\n| `Qwen3.8-35B-A3B-Q5_K_M.gguf` | Q5_K_M | 25.348 GB | Higher quality, modest size increase. |\n| `Qwen3.8-35B-A3B-Q6_K.gguf` | Q6_K | 29.209 GB | Near-lossless. |\n| `Qwen3.8-35B-A3B-Q8_…\n\nSource: https://huggingface.co/empero-ai/Qwen3.8-35B-A3B-Distill-GGUF","install":{"kind":"model","hfId":"empero-ai/Qwen3.8-35B-A3B-Distill-GGUF","gated":false,"format":"gguf","files":[{"name":"Qwen3.8-35B-A3B-BF16.gguf","size":71066994688,"quant":"BF16","sha256":"d5ff4a315a370b7b2ddaa7d0a756605f16daf8358f9e6173db08705f1fe0bb4f"},{"name":"Qwen3.8-35B-A3B-IQ2_M.gguf","size":12558245760,"quant":"IQ2_M","sha256":"9c095175f7af0c4acc18e552f0dd7ac8180d81f9ac4bf38d62a03c76a0d6e084"},{"name":"Qwen3.8-35B-A3B-IQ3_M.gguf","size":16339529600,"quant":"IQ3_M","sha256":"0ad7b253ecee6de7b38ecaf82ea5553d302c3f924961b7091b7088cc99a30f95"},{"name":"Qwen3.8-35B-A3B-IQ4_XS.gguf","size":19627788160,"quant":"IQ4_XS","sha256":"b645af45431ef9b41f43cae51c9323b2d0ca84f23031d483d2105c643ae58d65"},{"name":"Qwen3.8-35B-A3B-Q2_K.gguf","size":13838604160,"quant":"Q2_K","sha256":"a738c481e78099626bc43a9d4e5e0478266ab7bab5a788225da61310a1dd24dc"},{"name":"Qwen3.8-35B-A3B-Q3_K_M.gguf","size":17663774592,"quant":"Q3_K_M","sha256":"48a5197d37318de7984d0e1d38c66ae00ac5b873ea0920242a1889cc2d35c4a9"},{"name":"Qwen3.8-35B-A3B-Q4_K_M.gguf","size":21713462944,"quant":"Q4_K_M","sha256":"196103269085bc54c9b8f49ed21e9f53e1b56b465e8b796c6d8e31e06f63cfa5"},{"name":"Qwen3.8-35B-A3B-Q5_K_M.gguf","size":25347532448,"quant":"Q5_K_M","sha256":"f1903bac4ee3eec1f9013735298867c96d97dfb55c71f826a714b17214abc4ad"},{"name":"Qwen3.8-35B-A3B-Q6_K.gguf","size":29208731296,"quant":"Q6_K","sha256":"0bb743311ee5d58eeeeff67d42d0e62a52347539e825ceb320d1a2178003f47b"},{"name":"Qwen3.8-35B-A3B-Q8_0.gguf","size":37802149536,"quant":"Q8_0","sha256":"7d986d310e686a91cb514cdd819719b4f80ace899c9aaad7bfccfdf4670a9bb3"},{"name":"mmproj-Qwen3.8-35B-A3B-F16.gguf","size":899283520,"quant":"F16","sha256":"4381cb5110074396c2c7b39221fffae0c31886aa95b674de8e103d97edf58b94"}],"totalBytes":266066096704,"suggestedFile":"Qwen3.8-35B-A3B-Q4_K_M.gguf","requirements":{"ramGb":24,"diskBytes":21713462944,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp","ollama"],"command":"lsh models install hf:empero-ai/Qwen3.8-35B-A3B-Distill-GGUF"}}