LogiShell store Open app

Store › model › LLM

Qwen3.8-27B-NVFP4-MTP-GGUF

by esatapedico · source Hugging Face · updated 2026-08-22

apache-2.014 GB~17 GB RAMsource aliveunlabeled

A family of nine GGUF files of Qwen3.8-27B (the native vision-language 27B dense model, Gated DeltaNet + Gated Attention hybrid layout, 262,144-token native context, MTP speculative head), converted…

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-02 20:59 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
Qwen3.8-27B-NVFP4-MTP-COMPACT-LOW.gguf14 GB
Qwen3.8-27B-NVFP4-MTP-HIGH.gguf16 GB
Qwen3.8-27B-NVFP4-MTP-HIGHEST.gguf22 GB
Qwen3.8-27B-NVFP4-MTP-LOW.gguf14 GB
Qwen3.8-27B-NVFP4-MTP-MEDIUM.gguf15 GB
Qwen3.8-27B-NVFP4-MTP-MID-HIGH.gguf16 GB
Qwen3.8-27B-NVFP4-MTP-ORIG.gguf31 GB
Qwen3.8-27B-NVFP4-MTP-VERY-HIGH.gguf18 GB
Qwen3.8-27B-NVFP4-MTP-VERY-LOW.gguf14 GB
mmproj-BF16.ggufBF16888 MB

From the source README

Qwen3.8-27B-NVFP4-MTP-GGUF

A family of nine GGUF files of `Qwen3.8-27B` (the native vision-language 27B dense model, Gated DeltaNet + Gated Attention hybrid layout, 262,144-token native context, MTP speculative head), converted from unsloth/Qwen3.8-27B-NVFP4. The MTP (multi-token prediction) speculative head is baked into every file — no separate drafter needed.

  • `ORIG` — the source-preserving conversion: native NVFP4 MLP backbone + BF16 attention/embeddings (the source's F8 attention is dequantized to BF16 because GGML has no F8 tensor type). This is the largest file and the one all tiers are derived from.
  • `VERY-LOW` / `COMPACT-LOW` / `LOW` / `MEDIUM` / `MID-HIGH` / `HIGH` / `VERY-HIGH` — a compact family sharing a byte-identical 448-tensor NVFP4 backbone (all attention + MLP re-quantized to NVFP4), differing only in the 10 "extra" tensors (LM head, token embedding, MTP draft head). `COMPACT-LOW` fills the gap between `VERY-LOW` and `LOW` — slightly smaller than `LOW` while keeping a materially stronger LM head than `VERY-LOW` (Q4_K vs Q3_K). `MID-HIGH` sits between `MEDIUM` and `HIGH` with all three head groups at Q8_0.
  • `HIGHEST` — the top tier: keeps the source's native NVFP4 MLP (layers 0-55) exactly as in `ORIG`, restores Q8_0 for attention + the late MLP layers + the LM head, and keeps the token embedding + MTP head in BF16. The closest compact approximation of the source layout, for high-end GPUs.

The goal: keep native NVFP4 density across the whole model for Blackwell, and offer a size/precision ladder for the tensors that most affect output quality and decode speed. On our dual 16 GB Blackwell setup every tier fits and runs (see notes before treating any numbers as meaningful).

Vision works. The model is a native VLM (images and video). Pair any of these GGUFs with the…

Source: https://huggingface.co/esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF

Card id model:hf:esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF · collected 2026-10-02 20:59 UTC · JSON