LogiShell store Open app

Store › model › LLM

Sharp-Spark-X2.5-4B-GGUF

by peculiar-ragdoll · source Hugging Face · updated 2026-09-20

apache-2.03 MB~1 GB RAMsource aliveunlabeled

Dynamic imatrix GGUF quants of XHToken/Spark-X2.5-4B, a 4B dense long-context model, carrying our cyber-and-coding-weighted importance matrix, a heuristic per-tensor bit allocation, and the Sharp-Spa…

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-09 19:01 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
Sharp-Spark-X2.5-4B-Q4_K_XL.ggufQ4_K_XL2.5 GB
Sharp-Spark-X2.5-4B-Q5_K_XL.ggufQ5_K_XL3.0 GB
Sharp-Spark-X2.5-4B-Q6_K_XL.ggufQ6_K_XL3.4 GB
Spark-X2.5-4B-cyber.imatrix.gguf3 MB

From the source README

Sharp-Spark-X2.5-4B-GGUF

Dynamic imatrix GGUF quants of XHToken/Spark-X2.5-4B,
a 4B dense long-context model, carrying our cyber-and-coding-weighted importance matrix, a
heuristic per-tensor bit allocation, and the Sharp-Spark chat template.

This is a small, fast, long-context coder that fits 6 GB-VRAM GPUs and still holds usable speed out to six-figure context.

Who is this for?

This is the coder model for people with a small-VRAM GPU and ≤ 16GB RAM.

If you have more regular RAM than 16GB, try a (Cyber)TielCoder with partial GPU offloading instead. Its MoE architecture makes it fast even when it's split onto system RAM, and its about twice as capable.

Which quant should I download?

Grab the Q6_K_XL when you can: its performance is validated. The Q4 is a fallback.

| tier | size | pick it for |
|---|---:|---|
| Q4_K_XL | 2.67 GB | 4 GB VRAM GPU, or more context headroom |
| Q5_K_XL | 3.24 GB | 5 GB VRAM GPU, a step up from Q4 |
| Q6_K_XL | 3.61 GB | 6+ GB VRAM GPU, validated SWE performance |

If you find yourself needing a smaller quant than Q4, you should consider picking a natively smaller model instead. Small models like this are sensitive to quantization damage below Q6.

Long context

Spark is a hybrid-attention model: of its 36 layers, only 9 are full-attention (every 4th layer);
the other 27 use a 512-token sliding window. So the KV cache barely grows — only the 9 full layers
scale with context, which is what makes a 4B usable at six-figure context.

Card id model:hf:peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF · collected 2026-10-09 19:01 UTC · JSON