{"v":1,"id":"model:hf:CMSManhattan/JiRackUltra_1b","slug":"model-cmsmanhattan-jirackultra-1b","kind":"model","category":"llm","title":"JiRackUltra_1b","summary":"JiRack Ultra 1B (CPU) A fast and efficient ~1.5B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new Routing, Tool call, and Ro…","source":{"provider":"hf","ref":"CMSManhattan/JiRackUltra_1b","url":"https://huggingface.co/CMSManhattan/JiRackUltra_1b","rev":"18cbff46548f6f7acea6dc78d997514df50a7ec2","fetchedAt":"2026-10-02T20:59:11.664Z","etag":"W/\"2975-stbGUrdjqnhtp7fJZK7N/+OmnyU\""},"author":{"name":"CMSManhattan","url":"https://huggingface.co/CMSManhattan"},"license":{"spdx":"mit","raw":"mit","open":true},"metrics":{"downloads":2532192,"downloadsWeek":327838,"likes":0,"takenAt":"2026-10-02T20:59:11.664Z"},"tags":["safetensors","gguf","qwen2","text-generation","ternary","bitnet","1.58bit","cpu","qwen2.5","deepseek","efficient","low-memory","jirack","web-ui","routing","tool-call","robotics","conversational","en","zh","ja","ko","fr","es","pt","de","it","ru","ar","vi","th","endpoints_compatible"],"pipeline":"text-generation","links":{"github":"ggml-org/llama.cpp","npm":"node-llama-cpp"},"updatedAt":"2026-10-01T18:10:57.000Z","collectedAt":"2026-10-02T20:59:11.664Z","review":{"numbers":["2,532,192 downloads on Hugging Face","0 likes","license mit","1.0 GB for JiRackUltra_1b_Q4_K_M.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:59:11.664Z","http":200},"description":"# JiRack Ultra 1B (CPU)\nA fast and efficient ~1.5B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new **Routing**, **Tool call**, and **Robotics** tags. Built on a redesigned DeepSeek R1 architecture with native ternary (BitNet-style) support and ready-to-run GGUF quantizations.\n- JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert model in RAG deployments, with the ONNX JiRack Java server as an alternative.\n\n# JiRack Ternary Architedure & JiRack Tokenizer\n- Benefits high quality CPU inference TQ_2 on Llama.cpp and Ollama via QAT\n- Robotcs, Routing, Coding,  Multimedia, Advanced tool calling via CMSManhattan/JiRackPrecisionTokenizer\n\n## PARTNERSHIP\n\n- NVIDIA\n- FISERV\n\n# Ollama production support\n- We are working to support JiRack on Ollama for production systems also\n- added Jirack chat without reasoning feature  https://ollama.com/cmsmanhattan\n- Follow fresh Ollama platform updates\n\n# JiRack sevice options\n- Current quantizations were done from the FP16 model, but the model allows for more compression thanks to its ternary architecture.\n- If you need to do ternary compression, please write to me and I'll perform QAT from your dataset, tailored specifically to your task.\n- Plus double QAT via ONNX QAT.\n- Adapt train process to avoid catastrophic forgetting with NDA\n- Adapt train process to avoid fast plato in training with NDA\n- Convert model to TQ2_0 with support AVX2 and AVX-512 CPU instructions for high performance on CPU\n- QAT for TQ_2 Llama.cpp Ternarization docs https://huggingface.co/CMSManhattan/JiRackUltra_1b/blob/main/QAT_to_Llama.cpp_GGUF_TQ2_0_JirackUltra_1b.md\n- Adapts to agentic or instruct models for tool calling, using the JiRak tokenizer to enable high-quality tool calling on small models — built as a domain-specific tool expert.\n- Deployment and scale\n\n# JiRack Cod…\n\nSource: https://huggingface.co/CMSManhattan/JiRackUltra_1b","install":{"kind":"model","hfId":"CMSManhattan/JiRackUltra_1b","gated":false,"format":"gguf","files":[{"name":"JiRackUltra_1b.Q6_K.gguf","size":1464178592,"quant":"Q6_K","sha256":"8a7a9a4052ded2800d95b37a0bfc8b00630dc417b0e3410785203bab58238463"},{"name":"JiRackUltra_1b.Q8_0.gguf","size":1894532000,"quant":"Q8_0","sha256":"f56b2061f0f3bdae6de28da27cf0abebf13f36554fb1c0ba18414e08d88a25dd"},{"name":"JiRackUltra_1b.gguf","size":3560416160,"sha256":"19bb919b2deb433fba4b1116d857678def467b49ca881b5f9e6dc6ad27664744"},{"name":"JiRackUltra_1b_Q3_K_M.gguf","size":924455840,"quant":"Q3_K_M","sha256":"73cc639081a63a3793448991c1c1ca05ea82e5a9f65736b9e571f32aca10f31d"},{"name":"JiRackUltra_1b_Q4_K_M.gguf","size":1117320608,"quant":"Q4_K_M","sha256":"8db8cb25c578442e4a05a664e02c5ec90dda393671b9ef2ca02d37170bb0335f"},{"name":"model.pt","size":3554344056,"sha256":"1967c7a23020b5a895ba2469c6d3e1768eb127a8d59e29cb8671cf95345e248d"},{"name":"model.safetensors","size":3554214752,"sha256":"16bea0a35dba47a078ac1750e001512a02837da37c70c3c806244efbe9efefc8"}],"totalBytes":16069462008,"suggestedFile":"JiRackUltra_1b_Q4_K_M.gguf","requirements":{"ramGb":2,"diskBytes":1117320608,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp","ollama"],"command":"lsh models install hf:CMSManhattan/JiRackUltra_1b"}}