Store › model › LLM
Qwen3.8-27B-NVFP4-MTP-GGUF
by esatapedico · source Hugging Face · updated 2026-08-22
apache-2.014 GB~17 GB RAMsource aliveunlabeled
A family of nine GGUF files of Qwen3.8-27B (the native vision-language 27B dense model, Gated DeltaNet + Gated Attention hybrid layout, 262,144-token native context, MTP speculative head), converted…
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF
- License: apache-2.0
- Requirements: about 17 GB of RAM, 14 GB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp, ollama.
- Tags:
ggufnvfp4qwen3.8qwen3.5blackwellmtpspeculative-decodingvisionmultimodalllama.cpptext-generationenmultilingual
Numbers
- 752,373 downloads on Hugging Face
- 115 likes
- license apache-2.0
- 14 GB for Qwen3.8-27B-NVFP4-MTP-VERY-LOW.gguf
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:59 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
Qwen3.8-27B-NVFP4-MTP-COMPACT-LOW.gguf | 14 GB | |
Qwen3.8-27B-NVFP4-MTP-HIGH.gguf | 16 GB | |
Qwen3.8-27B-NVFP4-MTP-HIGHEST.gguf | 22 GB | |
Qwen3.8-27B-NVFP4-MTP-LOW.gguf | 14 GB | |
Qwen3.8-27B-NVFP4-MTP-MEDIUM.gguf | 15 GB | |
Qwen3.8-27B-NVFP4-MTP-MID-HIGH.gguf | 16 GB | |
Qwen3.8-27B-NVFP4-MTP-ORIG.gguf | 31 GB | |
Qwen3.8-27B-NVFP4-MTP-VERY-HIGH.gguf | 18 GB | |
Qwen3.8-27B-NVFP4-MTP-VERY-LOW.gguf | 14 GB | |
mmproj-BF16.gguf | BF16 | 888 MB |
From the source README
Qwen3.8-27B-NVFP4-MTP-GGUF
A family of nine GGUF files of `Qwen3.8-27B` (the native vision-language 27B dense model, Gated DeltaNet + Gated Attention hybrid layout, 262,144-token native context, MTP speculative head), converted from unsloth/Qwen3.8-27B-NVFP4. The MTP (multi-token prediction) speculative head is baked into every file — no separate drafter needed.
- `ORIG` — the source-preserving conversion: native NVFP4 MLP backbone + BF16 attention/embeddings (the source's F8 attention is dequantized to BF16 because GGML has no F8 tensor type). This is the largest file and the one all tiers are derived from.
- `VERY-LOW` / `COMPACT-LOW` / `LOW` / `MEDIUM` / `MID-HIGH` / `HIGH` / `VERY-HIGH` — a compact family sharing a byte-identical 448-tensor NVFP4 backbone (all attention + MLP re-quantized to NVFP4), differing only in the 10 "extra" tensors (LM head, token embedding, MTP draft head). `COMPACT-LOW` fills the gap between `VERY-LOW` and `LOW` — slightly smaller than `LOW` while keeping a materially stronger LM head than `VERY-LOW` (Q4_K vs Q3_K). `MID-HIGH` sits between `MEDIUM` and `HIGH` with all three head groups at Q8_0.
- `HIGHEST` — the top tier: keeps the source's native NVFP4 MLP (layers 0-55) exactly as in `ORIG`, restores Q8_0 for attention + the late MLP layers + the LM head, and keeps the token embedding + MTP head in BF16. The closest compact approximation of the source layout, for high-end GPUs.
The goal: keep native NVFP4 density across the whole model for Blackwell, and offer a size/precision ladder for the tensors that most affect output quality and decode speed. On our dual 16 GB Blackwell setup every tier fits and runs (see notes before treating any numbers as meaningful).
Vision works. The model is a native VLM (images and video). Pair any of these GGUFs with the…
Source: https://huggingface.co/esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF
Card id model:hf:esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF · collected 2026-10-02 20:59 UTC · JSON