{"v":1,"id":"model:hf:prism-ml/Ternary-Bonsai-2-27B-gguf","slug":"model-prism-ml-ternary-bonsai-2-27b-gguf","kind":"model","category":"llm","title":"Ternary-Bonsai-2-27B-gguf","summary":"Prism ML Website Whitepaper Demo & Examples Discord","source":{"provider":"hf","ref":"prism-ml/Ternary-Bonsai-2-27B-gguf","url":"https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf","rev":"b072e1d3b35a0a630cece372c2127528e0994386","fetchedAt":"2026-10-02T20:54:19.970Z","etag":"W/\"3258-b2sCsHWROsuctLkn6dFNyA/QeKk\""},"author":{"name":"prism-ml","url":"https://huggingface.co/prism-ml"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":3869715,"downloadsWeek":327838,"likes":2355,"stars":130156,"openIssues":2528,"lastRelease":{"tag":"v0.5.0","at":"2026-09-23T20:50:06Z"},"pushedAt":"2026-10-02T20:17:08Z","takenAt":"2026-10-02T20:54:19.970Z"},"tags":["llama.cpp","gguf","ternary","2-bit","llama-cpp","cuda","metal","on-device","hybrid-attention","prismml","bonsai","text-generation","endpoints_compatible","conversational"],"pipeline":"text-generation","links":{"github":"ggml-org/llama.cpp","npm":"node-llama-cpp"},"updatedAt":"2026-09-25T21:58:06.000Z","collectedAt":"2026-10-02T20:54:19.970Z","review":{"numbers":["3,869,715 downloads on Hugging Face","2,355 likes","license apache-2.0","50 GB for Ternary-Bonsai-2-27B-F16.gguf","130,156 stars on ggml-org/llama.cpp","2,528 open issues and PRs","last release v0.5.0 on 2026-09-23","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:54:19.970Z","http":200},"description":"Prism ML Website &nbsp;|&nbsp;\n  Whitepaper &nbsp;|&nbsp;\n  Demo &amp; Examples &nbsp;|&nbsp;\n  Discord\n\n# Bonsai 2 27B — GGUF\n\nFull 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU)\n\n> **\\~9.3x** smaller than FP16 (ideal) | **98.2%** of FP16 intelligence retained | **\\~47 tok/s** on an Apple M5 Max laptop\n\n## Highlights\n\n- **\\~5.9 GB** language model (down from \\~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU\n- **98.2% of FP16 intelligence retained**: 84.78 average across 14 thinking-mode benchmarks — far above the conventional IQ2_XXS build (72.59) at about 82% of its footprint, and within 0.4 points of UD-Q4_K_XL at three times the footprint\n- **Retains thinking, reasoning, and agentic behavior** deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within half a point of full precision (96.57), coding level with the baseline (89.42), agentic tool calling at 74.92\n- **End-to-end ternary language weights** across embeddings, attention projections, MLP projections, and LM head, at a *true* 1.72 bits per weight — no high-precision escape hatches behind a low-bit label; the vision tower ships as a separate Q8_0 mmproj pack\n- **262K-token context** on-device, kept practical by the Qwen3.8-27B hybrid-attention backbone (\\~75% linear attention)\n- **Two GGUF packings** with custom ternary hybrid-attention kernels for llama.cpp (CUDA, Metal) — **PTQ1_0** packs trits densely (1.75 bits/weight, 5.95 GB), **PQ2_0** stores each trit in a 2-bit slot (2.13 bits/weight, 7.21 GB); packed weights are consumed directly, never expanded back to FP16\n- **MLX companion**: also available as [Ternary-Bonsai-2-27B-mlx-2bit](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit) for native Apple Silicon inference\n\n## Resources\n\n- **[Whitepaper](https://github.com/PrismML-Eng/Bonsai-demo/blob/main/bonsai-2-27b-whitepape…\n\nSource: https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf","install":{"kind":"model","hfId":"prism-ml/Ternary-Bonsai-2-27B-gguf","gated":false,"format":"gguf","files":[{"name":"Ternary-Bonsai-2-27B-F16.gguf","size":53808408928,"quant":"F16","sha256":"f6f3b2c9b41956c34b379ec7c301dc936bc38d79b3c24c83388dd7d76000c180"},{"name":"Ternary-Bonsai-2-27B-PQ2_0.gguf","size":7206168928,"quant":"Q2_0","sha256":"3907dc1658db1f78a9826bf8d5bcb8dc65db0d466388937af57f2294fae62ec1"},{"name":"Ternary-Bonsai-2-27B-PTQ1_0.gguf","size":5946648928,"quant":"TQ1_0","sha256":"53107f530aa52eb00912263ab1ee29bd199261c87cd7b4ad4ca1318c1fe33ee3"},{"name":"Ternary-Bonsai-2-27B-mmproj-BF16.gguf","size":931145856,"quant":"BF16","sha256":"e287342d92332fa3577ed1d42e921dac9370c08da58ba9337fa450f6cc76cfd7"},{"name":"Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf","size":629246976,"quant":"Q8_0","sha256":"6807ede61d570bb86ba34b756a0fa109edc33668604de867c6ea6d8f1d631903"}],"totalBytes":68521619616,"suggestedFile":"Ternary-Bonsai-2-27B-F16.gguf","requirements":{"ramGb":59,"diskBytes":53808408928,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp","ollama"],"command":"lsh models install hf:prism-ml/Ternary-Bonsai-2-27B-gguf"}}