{"v":1,"id":"model:hf:prism-ml/Ternary-Bonsai-8B-gguf","slug":"model-prism-ml-ternary-bonsai-8b-gguf","kind":"model","category":"llm","title":"Ternary-Bonsai-8B-gguf","summary":"Prism ML Website White Paper Demo & Examples Discord","source":{"provider":"hf","ref":"prism-ml/Ternary-Bonsai-8B-gguf","url":"https://huggingface.co/prism-ml/Ternary-Bonsai-8B-gguf","rev":"c2aefbeb4b24469cd11579c3384b990404c17a30","fetchedAt":"2026-10-02T20:53:44.065Z","etag":"W/\"1d03-7dwKFcbwWzN28dse6QrcX2keL8k\""},"author":{"name":"prism-ml","url":"https://huggingface.co/prism-ml"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":371478,"downloadsWeek":327838,"likes":164,"stars":130156,"openIssues":2528,"lastRelease":{"tag":"v0.5.0","at":"2026-09-23T20:50:06Z"},"pushedAt":"2026-10-02T20:17:08Z","takenAt":"2026-10-02T20:53:44.065Z"},"tags":["gguf","ternary","1.58-bit","llama-cpp","q2_0","on-device","prismml","bonsai","text-generation","eval-results","endpoints_compatible","conversational"],"pipeline":"text-generation","links":{"github":"ggml-org/llama.cpp","npm":"node-llama-cpp"},"updatedAt":"2026-06-10T22:41:30.000Z","collectedAt":"2026-10-02T20:53:44.065Z","review":{"numbers":["371,478 downloads on Hugging Face","164 likes","license apache-2.0","15 GB for Ternary-Bonsai-8B-F16.gguf","130,156 stars on ggml-org/llama.cpp","2,528 open issues and PRs","last release v0.5.0 on 2026-09-23","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:53:44.065Z","http":200},"description":"Prism ML Website &nbsp;|&nbsp;\n  White Paper &nbsp;|&nbsp;\n  Demo &amp; Examples &nbsp;|&nbsp;\n  Discord\n\n# Ternary-Bonsai-8B-gguf\n\nTernary (1.58-bit) language model in GGUF Q2_0 format for `llama.cpp`\n\n  \n\n## Resources\n\n- **[White Paper](https://github.com/PrismML-Eng/Bonsai-demo/blob/main/ternary-bonsai-8b-whitepaper.pdf)**\n- **[Demo repo](https://github.com/PrismML-Eng/Bonsai-demo)** — examples for serving, benchmarking, and integrating Bonsai\n- **[Discord](https://discord.gg/prismml)** — community support and updates\n- **Kernels**: Q2_0 is not yet in mainline `llama.cpp`. Use our fork at [PrismML-Eng/llama.cpp](https://github.com/PrismML-Eng/llama.cpp) (`prism` branch, default) which adds Q2_0 support for CPU (NEON/generic) and Metal. Upstream PR coming soon.\n\n## Model Overview\n\n| Item             | Specification                                                            |\n| :--------------- | :----------------------------------------------------------------------- |\n| Base model       | Qwen3-8B                                                                 |\n| Parameters       | 8.19B (~6.95B non-embedding)                                             |\n| Architecture     | GQA (32 query / 8 KV heads), SwiGLU MLP, RoPE, RMSNorm                   |\n| Layers           | 36 Transformer decoder blocks                                            |\n| Context length   | 65,536 tokens                                                            |\n| Vocab size       | 151,936                                                                  |\n| Weight format    | GGUF Q2_0 g128: {-1, 0, +1} with FP16 group-wise scaling                 |\n| Packed Q2_0 size | **2.03 GiB** (2.18 GB)                                                   |\n| Ternary coverage | Embeddings, attention projections, MLP projections, LM head              |\n| License          | Apache 2.0…\n\nSource: https://huggingface.co/prism-ml/Ternary-Bonsai-8B-gguf","install":{"kind":"model","hfId":"prism-ml/Ternary-Bonsai-8B-gguf","gated":false,"format":"gguf","files":[{"name":"Ternary-Bonsai-8B-F16.gguf","size":16383663200,"quant":"F16","sha256":"a6abfaf896c1e36db825112fc0a18e49adea05eeca1c6b2fba4d785ca7e947ff"},{"name":"Ternary-Bonsai-8B-PQ2_0.gguf","size":2182184672,"quant":"Q2_0","sha256":"1376f942aa90e60f7b570c1d81b3916fea1315ff85aa1c4d19006af68fb4b922"},{"name":"Ternary-Bonsai-8B-Q2_0.gguf","size":2182184672,"quant":"Q2_0","sha256":"3c8d70470a5d97e5a2b9410ddd899cb740116591462626c60cb2fead6448f60b"},{"name":"Ternary-Bonsai-8B-Q2_0_g64.gguf","size":2310125920,"quant":"Q2_0","sha256":"e17b298d84ee78797916ae5c2ecc8211469cc65cccfe3080cd9a9bb503fbc55e"}],"totalBytes":23058158464,"suggestedFile":"Ternary-Bonsai-8B-F16.gguf","requirements":{"ramGb":19,"diskBytes":16383663200,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp","ollama"],"command":"lsh models install hf:prism-ml/Ternary-Bonsai-8B-gguf"}}