{"v":1,"id":"model:hf:prism-ml/Ternary-Bonsai-27B-gguf","slug":"model-prism-ml-ternary-bonsai-27b-gguf","kind":"model","category":"llm","title":"Ternary-Bonsai-27B-gguf","summary":"Prism ML Website Whitepaper Demo & Examples Discord","source":{"provider":"hf","ref":"prism-ml/Ternary-Bonsai-27B-gguf","url":"https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf","rev":"86e89f34c93201c3dfd5e5880fedb0022fc7e34d","fetchedAt":"2026-10-02T20:52:49.087Z","etag":"W/\"3060-46+OtwY9FOwscj2ZNpwXIjFZtOQ\""},"author":{"name":"prism-ml","url":"https://huggingface.co/prism-ml"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":628497,"downloadsWeek":327838,"likes":1403,"stars":130156,"openIssues":2528,"lastRelease":{"tag":"v0.5.0","at":"2026-09-23T20:50:06Z"},"pushedAt":"2026-10-02T20:17:08Z","takenAt":"2026-10-02T20:52:49.087Z"},"tags":["llama.cpp","gguf","conversational","ternary","2-bit","llama-cpp","cuda","metal","on-device","hybrid-attention","prismml","bonsai","text-generation","eval-results","endpoints_compatible"],"pipeline":"text-generation","links":{"github":"ggml-org/llama.cpp","npm":"node-llama-cpp"},"updatedAt":"2026-08-31T22:04:46.000Z","collectedAt":"2026-10-02T20:52:49.087Z","review":{"numbers":["628,497 downloads on Hugging Face","1,403 likes","license apache-2.0","50 GB for Ternary-Bonsai-27B-F16.gguf","130,156 stars on ggml-org/llama.cpp","2,528 open issues and PRs","last release v0.5.0 on 2026-09-23","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:52:49.087Z","http":200},"description":"Prism ML Website &nbsp;|&nbsp;\n  Whitepaper &nbsp;|&nbsp;\n  Demo &amp; Examples &nbsp;|&nbsp;\n  Discord\n\n# Ternary Bonsai 27B — GGUF\n\nFull 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU)\n\n> **\\~9.4x** smaller than FP16 (ideal) | **95%** of FP16 intelligence retained | **\\~26 tok/s** on an Apple M5 Pro laptop\n\n## Highlights\n\n- **\\~7.2 GB** deployed footprint (down from \\~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU\n- **95% of FP16 intelligence retained**: 80.49 average across 15 thinking-mode benchmarks — a *higher* score than the conventional IQ2_XXS build (72.73) at less than two-thirds of its footprint\n- **Retains thinking, reasoning, and agentic behavior** deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within two points of full precision (93.40), coding at 85.96, agentic tool use at 74.01\n- **End-to-end ternary language weights** across embeddings, attention projections, MLP projections, and LM head, at a *true* 1.71 bits per weight — no high-precision escape hatches behind a low-bit label; the vision tower ships in compact 4-bit HQQ\n- **262K-token context** on-device, kept practical by the Qwen3.6-27B hybrid-attention backbone (\\~75% linear attention) and 4-bit KV-cache quantization\n- **GGUF Q2_0_g128** format with custom 2-bit hybrid-attention kernels for llama.cpp (CUDA, Metal) — packed weights are consumed directly, never expanded back to FP16\n- **Ships with a DSpark speculative-decoding drafter layer** trained against the Bonsai 27B target — a lossless **1.34x** decode speedup on the CUDA serving path\n- **MLX companion**: also available as [Ternary-Bonsai-27B-mlx-2bit](https://huggingface.co/prism-ml/Ternary-Bonsai-27B-mlx-2bit) for native Apple Silicon inference\n- **1-bit companion**: the phone-class operating point (\\~3.9 GB) that fits an iPhone 17 Pro Max, published in GGUF as [Bonsa…\n\nSource: https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf","install":{"kind":"model","hfId":"prism-ml/Ternary-Bonsai-27B-gguf","gated":false,"format":"gguf","files":[{"name":"Ternary-Bonsai-27B-F16.gguf","size":53808280640,"quant":"F16","sha256":"f659ca3dd7e28ada5d8b5f3637862d0d51ef433bde032ec4c8990ed27c91a385"},{"name":"Ternary-Bonsai-27B-PQ2_0.gguf","size":7165121600,"quant":"Q2_0","sha256":"e4781999f1997ef97ce0c58d05750835acc999d18d83ee6489ba7ac7b14cb5f6"},{"name":"Ternary-Bonsai-27B-Q2_0.gguf","size":7165121600,"quant":"Q2_0","sha256":"868c11714cf8fe47f5ec9eeb2be0ab1a337112886f92ee0ede6b855c4fa31757"},{"name":"Ternary-Bonsai-27B-Q2_g64.gguf","size":7585330240,"quant":"Q2_G","sha256":"59a45d1ecef702b14531b06d22949f33b25c1897da31a8c0b298e01e4d9138eb"},{"name":"Ternary-Bonsai-27B-dspark-Q4_1.gguf","size":1946393568,"quant":"Q4_1","sha256":"c4810091d244eddc61a0cc4966e584b0959f141e3c66c0d371a6652d9f647da9"},{"name":"Ternary-Bonsai-27B-dspark-bf16.gguf","size":7291885792,"quant":"BF16","sha256":"d5ce05b0e7e23804279fb0b451e330e71c06d977657802a0e22d71433f30dbad"},{"name":"Ternary-Bonsai-27B-mmproj-BF16.gguf","size":931145760,"quant":"BF16","sha256":"acaf5b55d24ebd38c71fa220dc58c9a36776ec543b17728a4b322fc8d92f1de4"},{"name":"Ternary-Bonsai-27B-mmproj-Q8_0.gguf","size":629246880,"quant":"Q8_0","sha256":"eb561d41a7bbeb0fcf04883c8af11078ef6cae0a66862a0b68443cfca495269d"}],"totalBytes":86522526080,"suggestedFile":"Ternary-Bonsai-27B-F16.gguf","requirements":{"ramGb":59,"diskBytes":53808280640,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp","ollama"],"command":"lsh models install hf:prism-ml/Ternary-Bonsai-27B-gguf"}}