Store › model › LLM
deepseek-v4-gguf
by antirez · source Hugging Face · updated 2026-08-31
mit3.5 GB~5 GB RAMsource aliveunlabeled
This quants are specific for the DS4 inference engine. They may work with other inference engines or not (they should, but not the MTP model which requires a specific loader).
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:antirez/deepseek-v4-gguf and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/antirez/deepseek-v4-gguf
- License: mit
- Requirements: about 5 GB of RAM, 3.5 GB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp, ollama.
- Tags:
ggufquantizeddeepseekdeepseek-v4deepseek-v4-flashmoemixture-of-experts2-bit4-bitiq2_xxsq2_kq4_kds4apple-siliconmetaltext-generation
Numbers
- 1,487,199 downloads on Hugging Face
- 477 likes
- license mit
- 3.5 GB for DeepSeek-V4-Flash-MTP-Q4K-Q8_0-F32.gguf
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:59 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
DeepSeek-V4-Flash-DSpark-support-0731.gguf | 5.6 GB | |
DeepSeek-V4-Flash-DSpark-support.gguf | 5.6 GB | |
DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-0731.gguf | 81 GB | |
DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf | 81 GB | |
DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2.gguf | 81 GB | |
DeepSeek-V4-Flash-Layers37-42Q4KExperts-OtherExpertLayersIQ2XXSGateUp-Q2KDown-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-fixed-0731.gguf | 91 GB | |
DeepSeek-V4-Flash-Layers37-42Q4KExperts-OtherExpertLayersIQ2XXSGateUp-Q2KDown-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-fixed.gguf | 91 GB | |
DeepSeek-V4-Flash-MTP-Q4K-Q8_0-F32.gguf | Q8_0 | 3.5 GB |
DeepSeek-V4-Flash-MXFP4Experts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-mxfp4-0731.gguf | F16 | 145 GB |
DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix-0731.gguf | F16 | 153 GB |
DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-imatrix.gguf | F16 | 153 GB |
DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf | F16 | 153 GB |
DeepSeek-V4-Flash-Vision-Encoder.gguf | 890 MB | |
DeepSeek-V4-Flash-Vision-Exp-DSpark-support.gguf | 5.6 GB | |
DeepSeek-V4-Flash-Vision-Exp-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8.gguf | 81 GB | |
DeepSeek-V4-Flash-Vision-Exp-Layers37-42Q4KExperts-OtherExpertLayersIQ2XXSGateUp-Q2KDown-AProjQ8-SExpQ8-OutQ8.gguf | 91 GB | |
DeepSeek-V4-Flash-Vision-Exp-MXFP4Experts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out.gguf | F16 | 145 GB |
DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix-0813.gguf | 433 GB | |
DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix.gguf | 433 GB | |
DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct.gguf | 433 GB | |
DeepSeek-V4-Pro-Q4K-Layers-31-output.gguf | 412 GB | |
DeepSeek-V4-Pro-Q4K-Layers00-30.gguf | 426 GB |
From the source README
DeepSeek V4 Flash — GGUF for ds4
This quants are specific for the DS4 inference engine. They may work with other inference engines or not (they should, but not the MTP model which requires a specific loader).
https://github.com/antirez/ds4
Files
| File | Size | Routed experts (`ffn_{gate,up,down}_exps`) | Everything else |
|---|---:|---|---|
| `DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2.gguf` | 80.8 GiB | `IQ2_XXS` (gate, up) + `Q2_K` (down) | `Q8_0` attn proj / shared experts / output, `F16` router + embed + indexer + compressor + HC, `F32` norms / sinks / bias |
| `DeepSeek-V4-Flash-Q4KExperts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2.gguf` | 153.3 GiB | `Q4_K` (all three) | same as above |
| `DeepSeek-V4-Flash-MTP-Q4K-Q8_0-F32.gguf` | 3.6 GiB | MTP / speculative-decoding support (optional, not standalone). | |
Use q2 on 128 GB Mac machines, q4 on machines with ≥ 256 GB RAM, pair either with MTP for optional speculative decoding.
Quantization recipe
The filename is the spec. In detail, for the q2 file:
| Tensor class | Quant | Notes |
|---|---|---|
| `blk.*.ffn_gate_exps`, `blk.*.ffn_up_exps` | `IQ2_XXS` | routed-expert up/gate |
| `blk.*.ffn_down_exps` | `Q2_K` | routed-expert down (K-quant for quality) |
| `blk.*.ffn_{gate,up,down}_shexp` | `Q8_0` | shared experts |
| `blk.*.attn_q_a`, `attn_q_b`, `attn_kv`, `attn_output_a`, `attn_output_b` | `Q8_0` | all attention projections (MLA + low-rank output) |
| `output.weight` | `Q8_0` | output head |
| `token_embd.weight` | `F16` | input embedding |
| `blk.*.ffn_gate_inp` (router) | `F16` | learned router |
| `blk.*.exp_probs_b` (router bias), `blk.*.attn_sinks`, all `*_norm.weight` | `F32` | |
| `blk.*.ffn_gate_tid2eid` | `I32` | hash-routing tables (first 3 layers only) |
| `blk.*.attn_compressor_*`, `blk.*.indexer_*`, `blk.*.hc_*`, `blk.*.output_hc_*` | `F16` / `F32` | DSv4-specific…
Source: https://huggingface.co/antirez/deepseek-v4-gguf
Card id model:hf:antirez/deepseek-v4-gguf · collected 2026-10-02 20:59 UTC · JSON