{"v":1,"id":"model:hf:antirez/deepseek-v4-gguf","slug":"model-antirez-deepseek-v4-gguf","kind":"model","category":"llm","title":"deepseek-v4-gguf","summary":"This quants are specific for the DS4 inference engine. They may work with other inference engines or not (they should, but not the MTP model which requires a specific loader).","source":{"provider":"hf","ref":"antirez/deepseek-v4-gguf","url":"https://huggingface.co/antirez/deepseek-v4-gguf","rev":"f71f23d552d664e523b422157b2befbf74040380","fetchedAt":"2026-10-02T20:59:19.804Z","etag":"W/\"34e8-1NNv4bxFM/6LLpqsCwvxVHA1n60\""},"author":{"name":"antirez","url":"https://huggingface.co/antirez"},"license":{"spdx":"mit","raw":"mit","open":true},"metrics":{"downloads":1487199,"downloadsWeek":327838,"likes":477,"takenAt":"2026-10-02T20:59:19.804Z"},"tags":["gguf","quantized","deepseek","deepseek-v4","deepseek-v4-flash","moe","mixture-of-experts","2-bit","4-bit","iq2_xxs","q2_k","q4_k","ds4","apple-silicon","metal","text-generation","en","endpoints_compatible","conversational"],"pipeline":"text-generation","links":{"github":"deepseek-ai/DeepSeek-V3","npm":"node-llama-cpp"},"updatedAt":"2026-08-31T18:47:37.000Z","collectedAt":"2026-10-02T20:59:19.804Z","review":{"numbers":["1,487,199 downloads on Hugging Face","477 likes","license mit","3.5 GB for DeepSeek-V4-Flash-MTP-Q4K-Q8_0-F32.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:59:19.804Z","http":200},"description":"# DeepSeek V4 Flash — GGUF for ds4\n\nThis quants are specific for the DS4 inference engine. They may work with other inference engines or not (they should, but not the MTP model which requires a specific loader).\n\nhttps://github.com/antirez/ds4\n\n## Files\n\n| File | Size | Routed experts (`ffn_{gate,up,down}_exps`) | Everything else |\n|---|---:|---|---|\n| `DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2.gguf` | 80.8 GiB | `IQ2_XXS` (gate, up) + `Q2_K` (down) | `Q8_0` attn proj / shared experts / output, `F16` router + embed + indexer + compressor + HC, `F32` norms / sinks / bias |\n| `DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf` | 153.3 GiB | `Q4_K` (all three) | same as above |\n| `DeepSeek-V4-Flash-MTP-Q4K-Q8_0-F32.gguf` | 3.6 GiB | MTP / speculative-decoding support (optional, not standalone). | |\n\nUse **q2** on 128 GB Mac machines, **q4** on machines with ≥ 256 GB RAM, pair either with **MTP** for optional speculative decoding.\n\n## Quantization recipe\n\nThe filename is the spec. In detail, for the **q2** file:\n\n| Tensor class | Quant | Notes |\n|---|---|---|\n| `blk.*.ffn_gate_exps`, `blk.*.ffn_up_exps` | **`IQ2_XXS`** | routed-expert up/gate |\n| `blk.*.ffn_down_exps` | **`Q2_K`** | routed-expert down (K-quant for quality) |\n| `blk.*.ffn_{gate,up,down}_shexp` | `Q8_0` | shared experts |\n| `blk.*.attn_q_a`, `attn_q_b`, `attn_kv`, `attn_output_a`, `attn_output_b` | `Q8_0` | all attention projections (MLA + low-rank output) |\n| `output.weight` | `Q8_0` | output head |\n| `token_embd.weight` | `F16` | input embedding |\n| `blk.*.ffn_gate_inp` (router) | `F16` | learned router |\n| `blk.*.exp_probs_b` (router bias), `blk.*.attn_sinks`, all `*_norm.weight` | `F32` | |\n| `blk.*.ffn_gate_tid2eid` | `I32` | hash-routing tables (first 3 layers only) |\n| `blk.*.attn_compressor_*`, `blk.*.indexer_*`, `blk.*.hc_*`, `blk.*.output_hc_*` | `F16` / `F32` | DSv4-specific…\n\nSource: https://huggingface.co/antirez/deepseek-v4-gguf","install":{"kind":"model","hfId":"antirez/deepseek-v4-gguf","gated":false,"format":"gguf","files":[{"name":"DeepSeek-V4-Flash-DSpark-support-0731.gguf","size":5989114272,"sha256":"7e319924541db3f7a163ed7e11d7532a70d48228ab59d36cb81e1d4511885360"},{"name":"DeepSeek-V4-Flash-DSpark-support.gguf","size":5989114272,"sha256":"8b3adf5942bec22ae2ea867cd7079cf13530ba83ffcffaf00f5de48664a1a34e"},{"name":"DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-0731.gguf","size":86720111488,"sha256":"ca22ae2f838e14077c22bc1c1417b71b45b5e5a3687bd96c2ac6e17fdb6261c0"},{"name":"DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf","size":86720111488,"sha256":"efc7ed607ff27076e3e501fc3fefefa33c0ed8cf1eff483a2b7fdc0c2e616668"},{"name":"DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2.gguf","size":86720111200,"sha256":"31598c67c8b8744d3bcebcd19aa62253c6dc43cef3b8adf9f593656c9e86fd8c"},{"name":"DeepSeek-V4-Flash-Layers37-42Q4KExperts-OtherExpertLayersIQ2XXSGateUp-Q2KDown-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-fixed-0731.gguf","size":97591747456,"sha256":"659e22fbd01c9e13ea37a57c8d9c41e0a8819dffa3473d3c5286ee44b2d3398f"},{"name":"DeepSeek-V4-Flash-Layers37-42Q4KExperts-OtherExpertLayersIQ2XXSGateUp-Q2KDown-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-fixed.gguf","size":97591747456,"sha256":"edabc92af63ad8b139f00087fbfc10a4072f37b7597f4fd9ad1dfa6f83002396"},{"name":"DeepSeek-V4-Flash-MTP-Q4K-Q8_0-F32.gguf","size":3807602400,"quant":"Q8_0","sha256":"afd481ee689dce9037f70f39085fcdae5a5b096d521cdad43b19fa52bf8f4083"},{"name":"DeepSeek-V4-Flash-MXFP4Experts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-mxfp4-0731.gguf","size":155976458848,"quant":"F16","sha256":"0e3a161b670f686128ec5f92a601dfde616a37bf5e7e48999fa2d32471b57ec6"},{"name":"DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix-0731.gguf","size":164633502592,"quant":"F16","sha256":"6bb77b5ddcbc2d974c687cfb63d644ecfb295581b4a53fa4c1d810aea538254a"},{"name":"DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix.gguf","size":164633502592,"quant":"F16","sha256":"a2a3b31eca06344b93d32b2095511c4d36f92739a68a599b22047b4b2335d859"},{"name":"DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf","size":164633502304,"quant":"F16","sha256":"39e5de72ac544fdd5ffaf83ec28e36aaf3341b145235488e67d59400bbb3af55"},{"name":"DeepSeek-V4-Flash-Vision-Encoder.gguf","size":932857760,"sha256":"00cd4d81a435364967400a95c42703343e11da6b6f18c5143fe76e1d94d5035f"},{"name":"DeepSeek-V4-Flash-Vision-Exp-DSpark-support.gguf","size":5989114528,"sha256":"0807a67fd9ce5874bfc60d8d2461f50e11657e3dd94913d3473f85aa679bc877"},{"name":"DeepSeek-V4-Flash-Vision-Exp-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8.gguf","size":86720111776,"sha256":"8f2d42c0071ccf8a98f391cc2b835fd123f12330690b3059dbb7707920e5ad9e"},{"name":"DeepSeek-V4-Flash-Vision-Exp-Layers37-42Q4KExperts-OtherExpertLayersIQ2XXSGateUp-Q2KDown-AProjQ8-SExpQ8-OutQ8.gguf","size":97591747744,"sha256":"cded4517bb9d033e778e8bc4ccf1e79ba96d1c2d2b9f1c071c1d4a9037c51b02"},{"name":"DeepSeek-V4-Flash-Vision-Exp-MXFP4Experts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out.gguf","size":155976459136,"quant":"F16","sha256":"fc1efb96fa26e654b3530ce5f4b926b189a936d41d94dc1903c832f1e18eb3e7"},{"name":"DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix-0813.gguf","size":464627334560,"sha256":"c4d997ab9894b6c78b759f7869fe1726b6314b6515f6ff82607df3797c5eb193"},{"name":"DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf","size":464627334560,"sha256":"a0314d9c0e16122cd60071079124a2d17185d317c55a8f95ecb3ed3506278a96"},{"name":"DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct.gguf","size":464627334240,"sha256":"0e481c300b52414d9e415e42097eb69847616e8a02f8f457d97203e8d0ed60a5"},{"name":"DeepSeek-V4-Pro-Q4K-Layers-31-output.gguf","size":441962533120,"sha256":"41d14e4ccf9a9b777899887ac4d6115b11e5a5125f051e9fa5e727656ad5179b"},{"name":"DeepSeek-V4-Pro-Q4K-Layers00-30.gguf","size":457521327328,"sha256":"3c4526735ce204a99174059b216db155846b729bf5014c6b86d573323daa3cfa"}],"totalBytes":3761582781120,"suggestedFile":"DeepSeek-V4-Flash-MTP-Q4K-Q8_0-F32.gguf","requirements":{"ramGb":5,"diskBytes":3807602400,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp","ollama"],"command":"lsh models install hf:antirez/deepseek-v4-gguf"}}