{"v":1,"id":"model:hf:michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF","slug":"model-michaelw9999-qwen3-6-35b-a3b-nvfp4-mtp-gguf","kind":"model","category":"llm","title":"Qwen3.6-35B-A3B-NVFP4-MTP-GGUF","summary":"This repo contains two experimental NVFP4 GGUF quantizations of Qwen3.6-35B-A3B for llama.cpp. This was quantized using my experimental advanced-gguf-quantizer tool. Both models were imatrix calibrat…","source":{"provider":"hf","ref":"michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF","url":"https://huggingface.co/michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF","rev":"df112dd576e55b1daa1331a7831b64ec9c03dbae","fetchedAt":"2026-10-02T21:00:53.533Z","etag":"W/\"27a3-Env9GzCJY/Ni0al8VOacy8A0VVE\""},"author":{"name":"michaelw9999","url":"https://huggingface.co/michaelw9999"},"license":{"spdx":null,"raw":null,"open":null,"note":"the source did not name a license"},"metrics":{"downloads":365040,"downloadsWeek":327838,"likes":12,"takenAt":"2026-10-02T21:00:53.533Z"},"tags":["gguf","qwen3.6","qwen3.6-35b","nvfp4","llama.cpp","michaelw9999","qwen","blackwell","text-generation","endpoints_compatible","imatrix","conversational"],"pipeline":"text-generation","links":{"github":"QwenLM/Qwen3","npm":"node-llama-cpp"},"updatedAt":"2026-06-12T08:37:03.000Z","collectedAt":"2026-10-02T21:00:53.533Z","review":{"numbers":["365,040 downloads on Hugging Face","12 likes","license not named","19 GB for Qwen3.6-35B-A3B-NVFP4-MTP-TURBO.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T21:00:53.533Z","http":200},"description":"# Qwen3.6-35B-A3B-NVFP4-MTP-GGUF\n\nThis repo contains two experimental NVFP4 GGUF quantizations of **Qwen3.6-35B-A3B** for `llama.cpp`.\nThis was quantized using my experimental advanced-gguf-quantizer tool.\nBoth models were imatrix calibrated for the first time using a new custom dataset that I am evaluating.\n\nThis repository contains two NVFP4 variants:\n\n| Variant | File | Best for | Notes |\n|---|---|---|---|\n| **TURBO** | [`Qwen3.6-35B-A3B-NVFP4-MTP-TURBO.gguf`](./Qwen3.6-35B-A3B-NVFP4-MTP-TURBO.gguf) | Max  speed | More NVFP4. Lower quality metrics. |\n| **HQ** | [`Qwen3.6-35B-A3B-NVFP4-MTP-HQ.gguf`](./Qwen3.6-35B-A3B-NVFP4-MTP-HQ.gguf) | Better quality | More tensors promoted. Slightly slower. |\n\n## Quality & Speed Results\n\nAll PPL/KLD results were measured against the same BF16 wikitest KLD base, and then compared to the official NVFP4 release by NVIDIA.\n\n| Metric | TURBO | HQ |  NVIDIA-NVFP4 |\n|---|---:|---:|---:|\n| Size | **18.56 GiB** | 18.64 GiB | 22.20 GiB |\n| Mean PPL(Q) | 6.987392 | **6.897796** | 7.014030 |\n| Mean PPL(Q)-PPL(base) | 0.268551 | **0.178955** | — |\n| Mean PPL ratio | 1.039970 | **1.026635** | 1.043935 |\n| Mean ln(PPL ratio) | 0.039192 | **0.026286** | — |\n| Mean KLD | 0.063228 | **0.050759** | 0.066331 |\n| 99.9% KLD | 1.924147 | 1.565143 | **1.560988** |\n| 99.0% KLD | 0.598519 | **0.488387** | 0.495896 |\n| 95.0% KLD | 0.221030 | **0.178889** | 0.207580 |\n| Max KLD | 11.946571 | 10.093911 | **6.972712** |\n| Same top p | 89.023% | **90.255%** | 87.608% |\n| Top flip weight | 0.012068 | **0.009575** | — |\n| pp512 | **11593.57 t/s** | 10936.20 t/s | 10426.32 t/s |\n| tg128 | **271.21 t/s** | 270.49 t/s | 221.86 t/s |\n\n## Evaluation Results\n\nFurther evaluation tests are underway to identify real world performance differences between **TURBO** and **HQ**.\n\n| Benchmark | Samples |   TURBO   |   HQ   | NVIDIA-NVFP4 |\n|---|---:|---:|---:|---:|\n| GSM8K | 103 | 98% | 98% | 97% |…\n\nSource: https://huggingface.co/michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF","install":{"kind":"model","hfId":"michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF","gated":false,"format":"gguf","files":[{"name":"Qwen3.6-35B-A3B-NVFP4-MTP-HQ.gguf","size":20487740864,"sha256":"777564174a7ccf01a2e9d171ac73206ec3da6b6f6b0124e71a9628ac19f61aa9"},{"name":"Qwen3.6-35B-A3B-NVFP4-MTP-TURBO.gguf","size":20407437600,"sha256":"f3d2fdc74e3ef19925ccbf794b04d7f6f11fb12eba7722b7749219d0cc5c36ed"}],"totalBytes":40895178464,"suggestedFile":"Qwen3.6-35B-A3B-NVFP4-MTP-TURBO.gguf","requirements":{"ramGb":23,"diskBytes":20407437600,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp","ollama"],"command":"lsh models install hf:michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF"}}