{"v":1,"id":"model:hf:esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF","slug":"model-esatapedico-qwen3-8-27b-nvfp4-mtp-gguf","kind":"model","category":"llm","title":"Qwen3.8-27B-NVFP4-MTP-GGUF","summary":"A family of nine GGUF files of Qwen3.8-27B (the native vision-language 27B dense model, Gated DeltaNet + Gated Attention hybrid layout, 262,144-token native context, MTP speculative head), converted…","source":{"provider":"hf","ref":"esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF","url":"https://huggingface.co/esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF","rev":"bcd7a7d3e251d4ec0fd15c72584b5eb9e0981383","fetchedAt":"2026-10-02T20:59:37.626Z","etag":"W/\"1440-B7NDxBfGxtt0KImhHBptHpugPOE\""},"author":{"name":"esatapedico","url":"https://huggingface.co/esatapedico"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":752373,"downloadsWeek":327838,"likes":115,"takenAt":"2026-10-02T20:59:37.626Z"},"tags":["gguf","nvfp4","qwen3.8","qwen3.5","blackwell","mtp","speculative-decoding","vision","multimodal","llama.cpp","text-generation","en","multilingual"],"pipeline":"text-generation","links":{"github":"QwenLM/Qwen3","npm":"node-llama-cpp"},"updatedAt":"2026-08-22T00:11:17.000Z","collectedAt":"2026-10-02T20:59:37.626Z","review":{"numbers":["752,373 downloads on Hugging Face","115 likes","license apache-2.0","14 GB for Qwen3.8-27B-NVFP4-MTP-VERY-LOW.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:59:37.626Z","http":200},"description":"# Qwen3.8-27B-NVFP4-MTP-GGUF\n\nA **family of nine GGUF files** of `Qwen3.8-27B` (the native vision-language 27B dense model, Gated DeltaNet + Gated Attention hybrid layout, 262,144-token native context, MTP speculative head), converted from [unsloth/Qwen3.8-27B-NVFP4](https://huggingface.co/unsloth/Qwen3.8-27B-NVFP4). The **MTP (multi-token prediction) speculative head is baked into every file** — no separate drafter needed.\n\n- **`ORIG`** — the source-preserving conversion: native **NVFP4 MLP backbone** + **BF16 attention/embeddings** (the source's F8 attention is dequantized to BF16 because GGML has no F8 tensor type). This is the largest file and the one all tiers are derived from.\n- **`VERY-LOW` / `COMPACT-LOW` / `LOW` / `MEDIUM` / `MID-HIGH` / `HIGH` / `VERY-HIGH`** — a compact family sharing a **byte-identical 448-tensor NVFP4 backbone** (all attention + MLP re-quantized to NVFP4), differing only in the 10 \"extra\" tensors (LM head, token embedding, MTP draft head). `COMPACT-LOW` fills the gap between `VERY-LOW` and `LOW` — slightly smaller than `LOW` while keeping a materially stronger LM head than `VERY-LOW` (Q4_K vs Q3_K). `MID-HIGH` sits between `MEDIUM` and `HIGH` with all three head groups at **Q8_0**.\n- **`HIGHEST`** — the top tier: keeps the source's native **NVFP4 MLP** (layers 0-55) exactly as in `ORIG`, restores **Q8_0** for attention + the late MLP layers + the LM head, and keeps the **token embedding + MTP head in BF16**. The closest compact approximation of the source layout, for high-end GPUs.\n\nThe goal: **keep native NVFP4 density across the whole model for Blackwell**, and offer a size/precision ladder for the tensors that most affect output quality and decode speed. On our dual 16 GB Blackwell setup every tier fits and runs (see notes before treating any numbers as meaningful).\n\n**Vision works.** The model is a native VLM (images and video). Pair any of these GGUFs with the…\n\nSource: https://huggingface.co/esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF","install":{"kind":"model","hfId":"esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF","gated":false,"format":"gguf","files":[{"name":"Qwen3.8-27B-NVFP4-MTP-COMPACT-LOW.gguf","size":15160261920,"sha256":"ac0ef9c5eceb5a5dc9b266eacc9372508158713c6c35618cf735675451fdd3ac"},{"name":"Qwen3.8-27B-NVFP4-MTP-HIGH.gguf","size":17570799040,"sha256":"d57008707b0558bde05ce61d7402e4e668ffd45a4c97d03d2ad97db73f98d403"},{"name":"Qwen3.8-27B-NVFP4-MTP-HIGHEST.gguf","size":23185001824,"sha256":"6a202c2faf67f79d4c8c61ec940da7a62bd59a87608508fe8048131630cc4ba6"},{"name":"Qwen3.8-27B-NVFP4-MTP-LOW.gguf","size":15534575072,"sha256":"ce66a629d4a3516bba27ca91de29372f086f90f72ddb92fe298de67b8bb88bbc"},{"name":"Qwen3.8-27B-NVFP4-MTP-MEDIUM.gguf","size":16378863040,"sha256":"f0b4c538c75037f026bde3b650f0ca639d382c128a4572769dce1183db86253a"},{"name":"Qwen3.8-27B-NVFP4-MTP-MID-HIGH.gguf","size":16912387392,"sha256":"79b032f7a118fb34f1445c4d7ae50bc7b304c74161035c3df8bd5526c12899b9"},{"name":"Qwen3.8-27B-NVFP4-MTP-ORIG.gguf","size":33133121024,"sha256":"f18098dfc32ca398f63093c113d48bd962684c78a3e93d4158f8f440699f7cae"},{"name":"Qwen3.8-27B-NVFP4-MTP-VERY-HIGH.gguf","size":19694390752,"sha256":"3e52d6280ee650520a2d901002121c11cf9d23ac75f23c52a362bf285d561d81"},{"name":"Qwen3.8-27B-NVFP4-MTP-VERY-LOW.gguf","size":14862277984,"sha256":"74ea17ea05e0e0241af8d5b29cdea38b3f4509f66d9b96c1ab05f0e1f0e537d9"},{"name":"mmproj-BF16.gguf","size":931146432,"quant":"BF16","sha256":"83ee4f4f205fa514161778c41df1ea14144faa0f713510893b63c2395f5c2d53"}],"totalBytes":173362824480,"suggestedFile":"Qwen3.8-27B-NVFP4-MTP-VERY-LOW.gguf","requirements":{"ramGb":17,"diskBytes":14862277984,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp","ollama"],"command":"lsh models install hf:esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF"}}