Store › model › LLM
Qwen3.6-35B-A3B-NVFP4-MTP-GGUF
by michaelw9999 · source Hugging Face · updated 2026-06-12
license unknown19 GB~23 GB RAMsource aliveunlabeled
This repo contains two experimental NVFP4 GGUF quantizations of Qwen3.6-35B-A3B for llama.cpp. This was quantized using my experimental advanced-gguf-quantizer tool. Both models were imatrix calibrat…
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF
- License: not named by the source (the source did not name a license)
- Requirements: about 23 GB of RAM, 19 GB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp, ollama.
- Tags:
ggufqwen3.6qwen3.6-35bnvfp4llama.cppmichaelw9999qwenblackwelltext-generationendpoints_compatibleimatrixconversational
Numbers
- 365,040 downloads on Hugging Face
- 12 likes
- license not named
- 19 GB for Qwen3.6-35B-A3B-NVFP4-MTP-TURBO.gguf
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 21:00 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
Qwen3.6-35B-A3B-NVFP4-MTP-HQ.gguf | 19 GB | |
Qwen3.6-35B-A3B-NVFP4-MTP-TURBO.gguf | 19 GB |
From the source README
Qwen3.6-35B-A3B-NVFP4-MTP-GGUF
This repo contains two experimental NVFP4 GGUF quantizations of Qwen3.6-35B-A3B for `llama.cpp`.
This was quantized using my experimental advanced-gguf-quantizer tool.
Both models were imatrix calibrated for the first time using a new custom dataset that I am evaluating.
This repository contains two NVFP4 variants:
| Variant | File | Best for | Notes |
|---|---|---|---|
| TURBO | `Qwen3.6-35B-A3B-NVFP4-MTP-TURBO.gguf` | Max speed | More NVFP4. Lower quality metrics. |
| HQ | `Qwen3.6-35B-A3B-NVFP4-MTP-HQ.gguf` | Better quality | More tensors promoted. Slightly slower. |
Quality & Speed Results
All PPL/KLD results were measured against the same BF16 wikitest KLD base, and then compared to the official NVFP4 release by NVIDIA.
| Metric | TURBO | HQ | NVIDIA-NVFP4 |
|---|---:|---:|---:|
| Size | 18.56 GiB | 18.64 GiB | 22.20 GiB |
| Mean PPL(Q) | 6.987392 | 6.897796 | 7.014030 |
| Mean PPL(Q)-PPL(base) | 0.268551 | 0.178955 | — |
| Mean PPL ratio | 1.039970 | 1.026635 | 1.043935 |
| Mean ln(PPL ratio) | 0.039192 | 0.026286 | — |
| Mean KLD | 0.063228 | 0.050759 | 0.066331 |
| 99.9% KLD | 1.924147 | 1.565143 | 1.560988 |
| 99.0% KLD | 0.598519 | 0.488387 | 0.495896 |
| 95.0% KLD | 0.221030 | 0.178889 | 0.207580 |
| Max KLD | 11.946571 | 10.093911 | 6.972712 |
| Same top p | 89.023% | 90.255% | 87.608% |
| Top flip weight | 0.012068 | 0.009575 | — |
| pp512 | 11593.57 t/s | 10936.20 t/s | 10426.32 t/s |
| tg128 | 271.21 t/s | 270.49 t/s | 221.86 t/s |
Evaluation Results
Further evaluation tests are underway to identify real world performance differences between TURBO and HQ.
| Benchmark | Samples | TURBO | HQ | NVIDIA-NVFP4 |
|---|---:|---:|---:|---:|
| GSM8K | 103 | 98% | 98% | 97% |…
Source: https://huggingface.co/michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF
Card id model:hf:michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF · collected 2026-10-02 21:00 UTC · JSON