{"v":1,"id":"model:hf:peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF","slug":"model-peculiar-ragdoll-sharp-spark-x2-5-4b-gguf","kind":"model","category":"llm","title":"Sharp-Spark-X2.5-4B-GGUF","summary":"Dynamic imatrix GGUF quants of XHToken/Spark-X2.5-4B, a 4B dense long-context model, carrying our cyber-and-coding-weighted importance matrix, a heuristic per-tensor bit allocation, and the Sharp-Spa…","source":{"provider":"hf","ref":"peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF","url":"https://huggingface.co/peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF","rev":"e797ddf6a57d9ecfddf68394438d2667ecb42dad","fetchedAt":"2026-10-09T19:01:12.056Z","etag":"W/\"5c26-o+5Twar6uKpRiSY90sANXcT0ZEk\""},"author":{"name":"peculiar-ragdoll","url":"https://huggingface.co/peculiar-ragdoll"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":326908,"downloadsWeek":297971,"likes":118,"stars":130649,"openIssues":2491,"lastRelease":{"tag":"v0.6.0","at":"2026-10-05T16:56:22Z"},"pushedAt":"2026-10-09T18:49:42Z","takenAt":"2026-10-09T19:01:12.056Z"},"tags":["gguf","llama.cpp","imatrix","spark2_5","long-context","coding","sharp-template","token-efficient","efficient-thinking","text-generation","conversational","endpoints_compatible"],"pipeline":"text-generation","links":{"github":"ggml-org/llama.cpp","npm":"node-llama-cpp"},"updatedAt":"2026-09-20T21:03:52.000Z","collectedAt":"2026-10-09T19:01:12.056Z","review":{"numbers":["326,908 downloads on Hugging Face","118 likes","license apache-2.0","0.0 GB for Spark-X2.5-4B-cyber.imatrix.gguf","130,649 stars on ggml-org/llama.cpp","2,491 open issues and PRs","last release v0.6.0 on 2026-10-05","297,971 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-09T19:01:12.056Z","http":200},"description":"# Sharp-Spark-X2.5-4B-GGUF\n\nDynamic **imatrix** GGUF quants of [XHToken/Spark-X2.5-4B](https://huggingface.co/XHToken/Spark-X2.5-4B),\na 4B dense long-context model, carrying our cyber-and-coding-weighted importance matrix, a\nheuristic per-tensor bit allocation, and the **Sharp-Spark** chat template.\n\nThis is a **small, fast, long-context coder that fits 6 GB-VRAM GPUs** and still holds usable speed out to six-figure context.\n\n---\n\n## Who is this for?\n\nThis is the coder model for people with a small-VRAM GPU and ≤ 16GB RAM.\n\nIf you have more regular RAM than 16GB, try a ([Cyber](https://huggingface.co/collections/peculiar-ragdoll/cyber-tiel-coder-35b-a3b))[TielCoder](https://huggingface.co/collections/peculiar-ragdoll/tiel-coder-35b-a3b) with partial GPU offloading instead. Its MoE architecture makes it fast even when it's split onto system RAM, and its about twice as capable.\n\n---\n\n## Which quant should I download?\n\n**Grab the Q6_K_XL** when you can: its performance is validated. The Q4 is a fallback.\n\n| tier | size | pick it for |\n|---|---:|---|\n| **Q4_K_XL** | 2.67 GB | 4 GB VRAM GPU, or more context headroom |\n| **Q5_K_XL** | 3.24 GB | 5 GB VRAM GPU, a step up from Q4 |\n| **Q6_K_XL** | 3.61 GB | 6+ GB VRAM GPU, validated SWE performance |\n\nIf you find yourself needing a smaller quant than Q4, you should consider picking a natively smaller model instead. Small models like this are sensitive to quantization damage below Q6.\n\n---\n\n## Long context\n\nSpark is a hybrid-attention model: of its 36 layers, **only 9 are full-attention** (every 4th layer);\nthe other 27 use a **512-token sliding window**. So the KV cache barely grows — only the 9 full layers\nscale with context, which is what makes a 4B usable at six-figure context.\n\n**Short vs 131k context:**\n\n| context | prefill tok/s | decode tok/s |\n|---|---:|---:|\n| short (512) | 1684 | 111 |\n| **131,072** | **557** | **42** |\n\nDecode fell only **2…\n\nSource: https://huggingface.co/peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF","install":{"kind":"model","hfId":"peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF","gated":false,"format":"gguf","files":[{"name":"Sharp-Spark-X2.5-4B-Q4_K_XL.gguf","size":2677961984,"quant":"Q4_K_XL","sha256":"8e5601dbd18fbc2b731cf674a040dd32f3ec2d09a312f4e0f3c4d7bc92998837"},{"name":"Sharp-Spark-X2.5-4B-Q5_K_XL.gguf","size":3246363904,"quant":"Q5_K_XL","sha256":"f445f1a57e58b70ea85078e1edcd29763843f71f154bac2efc57eea1b8333a26"},{"name":"Sharp-Spark-X2.5-4B-Q6_K_XL.gguf","size":3612423424,"quant":"Q6_K_XL","sha256":"793e673f34d2dde9674d24d277c25dbf03b89290333835aa31b7ee1d62e20dfc"},{"name":"Spark-X2.5-4B-cyber.imatrix.gguf","size":3572704,"sha256":"c0ccf6c2e49d7565db46dc6eb026b0b1e01e747d103bd04fd9e9cfe5a2f64a97"}],"totalBytes":9540322016,"suggestedFile":"Spark-X2.5-4B-cyber.imatrix.gguf","requirements":{"ramGb":1,"diskBytes":3572704,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp","ollama"],"command":"lsh models install hf:peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF"}}