Store › model › LLM
Sharp-Spark-X2.5-4B-GGUF
by peculiar-ragdoll · source Hugging Face · updated 2026-09-20
apache-2.03 MB~1 GB RAMsource aliveunlabeled
Dynamic imatrix GGUF quants of XHToken/Spark-X2.5-4B, a 4B dense long-context model, carrying our cyber-and-coding-weighted importance matrix, a heuristic per-tensor bit allocation, and the Sharp-Spa…
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF
- License: apache-2.0
- Requirements: about 1 GB of RAM, 3 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp, ollama.
- Tags:
ggufllama.cppimatrixspark2_5long-contextcodingsharp-templatetoken-efficientefficient-thinkingtext-generationconversationalendpoints_compatible
Numbers
- 326,908 downloads on Hugging Face
- 118 likes
- license apache-2.0
- 0.0 GB for Spark-X2.5-4B-cyber.imatrix.gguf
- 130,649 stars on ggml-org/llama.cpp
- 2,491 open issues and PRs
- last release v0.6.0 on 2026-10-05
- 297,971 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-09 19:01 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
Sharp-Spark-X2.5-4B-Q4_K_XL.gguf | Q4_K_XL | 2.5 GB |
Sharp-Spark-X2.5-4B-Q5_K_XL.gguf | Q5_K_XL | 3.0 GB |
Sharp-Spark-X2.5-4B-Q6_K_XL.gguf | Q6_K_XL | 3.4 GB |
Spark-X2.5-4B-cyber.imatrix.gguf | 3 MB |
From the source README
Sharp-Spark-X2.5-4B-GGUF
Dynamic imatrix GGUF quants of XHToken/Spark-X2.5-4B,
a 4B dense long-context model, carrying our cyber-and-coding-weighted importance matrix, a
heuristic per-tensor bit allocation, and the Sharp-Spark chat template.
This is a small, fast, long-context coder that fits 6 GB-VRAM GPUs and still holds usable speed out to six-figure context.
Who is this for?
This is the coder model for people with a small-VRAM GPU and ≤ 16GB RAM.
If you have more regular RAM than 16GB, try a (Cyber)TielCoder with partial GPU offloading instead. Its MoE architecture makes it fast even when it's split onto system RAM, and its about twice as capable.
Which quant should I download?
Grab the Q6_K_XL when you can: its performance is validated. The Q4 is a fallback.
| tier | size | pick it for |
|---|---:|---|
| Q4_K_XL | 2.67 GB | 4 GB VRAM GPU, or more context headroom |
| Q5_K_XL | 3.24 GB | 5 GB VRAM GPU, a step up from Q4 |
| Q6_K_XL | 3.61 GB | 6+ GB VRAM GPU, validated SWE performance |
If you find yourself needing a smaller quant than Q4, you should consider picking a natively smaller model instead. Small models like this are sensitive to quantization damage below Q6.
Long context
Spark is a hybrid-attention model: of its 36 layers, only 9 are full-attention (every 4th layer);
the other 27 use a 512-token sliding window. So the KV cache barely grows — only the 9 full layers
scale with context, which is what makes a 4B usable at six-figure context.
Card id model:hf:peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF · collected 2026-10-09 19:01 UTC · JSON