{"v":1,"id":"model:hf:prism-ml/Bonsai-27B-gguf","slug":"model-prism-ml-bonsai-27b-gguf","kind":"model","category":"llm","title":"Bonsai-27B-gguf","summary":"Prism ML Website Whitepaper Demo & Examples Discord","source":{"provider":"hf","ref":"prism-ml/Bonsai-27B-gguf","url":"https://huggingface.co/prism-ml/Bonsai-27B-gguf","rev":"f10afb355f104535e3e3e98cf7ab7795c72bd292","fetchedAt":"2026-10-02T20:53:35.361Z","etag":"W/\"3004-YEqegeXMqJdPtJQEhFgzZilSNKw\""},"author":{"name":"prism-ml","url":"https://huggingface.co/prism-ml"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":425798,"downloadsWeek":327838,"likes":883,"stars":130156,"openIssues":2528,"lastRelease":{"tag":"v0.5.0","at":"2026-09-23T20:50:06Z"},"pushedAt":"2026-10-02T20:17:08Z","takenAt":"2026-10-02T20:53:35.361Z"},"tags":["llama.cpp","gguf","conversational","1-bit","llama-cpp","cuda","metal","on-device","hybrid-attention","prismml","bonsai","text-generation","eval-results","endpoints_compatible"],"pipeline":"text-generation","links":{"github":"ggml-org/llama.cpp","npm":"node-llama-cpp"},"updatedAt":"2026-07-17T16:05:42.000Z","collectedAt":"2026-10-02T20:53:35.361Z","review":{"numbers":["425,798 downloads on Hugging Face","883 likes","license apache-2.0","50 GB for Bonsai-27B-F16.gguf","130,156 stars on ggml-org/llama.cpp","2,528 open issues and PRs","last release v0.5.0 on 2026-09-23","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:53:35.361Z","http":200},"description":"Prism ML Website &nbsp;|&nbsp;\n  Whitepaper &nbsp;|&nbsp;\n  Demo &amp; Examples &nbsp;|&nbsp;\n  Discord\n\n# 1-bit Bonsai 27B — GGUF\n\nFull 27B-class reasoning in binary transformer weights, for llama.cpp (CUDA, Metal, CPU)\n\n> **\\~14.2x** smaller than FP16 | **\\~90%** of FP16 intelligence retained | **\\~44 tok/s** on an Apple M5 Pro laptop\n\n## Highlights\n\n- **\\~3.9 GB** deployed footprint (down from \\~54 GB FP16) — a 27B model on everyday laptops and single GPUs\n- **Retains thinking, reasoning, and agentic behavior** deep in the sub-4-bit regime, where conventional low-bit representations collapse — 76.11 average across 15 thinking-mode benchmarks (89.5% of FP16), including math at 91.66 and coding at 81.88\n- **End-to-end binary language weights** across embeddings, attention projections, MLP projections, and LM head, at a *true* 1.125 bits per weight — no high-precision escape hatches behind a low-bit label; the vision tower ships in compact 4-bit HQQ\n- **262K-token context** on-device, kept practical by the Qwen3.6-27B hybrid-attention backbone (\\~75% linear attention) and 4-bit KV-cache quantization\n- **GGUF Q1_0_g128** format with custom 1-bit hybrid-attention kernels for llama.cpp (CUDA, Metal) — packed weights are consumed directly, never expanded back to FP16\n- **Ships with a DSpark speculative-decoding drafter layer** trained against the Bonsai 27B target — a lossless **1.37x** decode speedup on the CUDA serving path\n- **MLX companion**: also available as [Bonsai-27B-mlx-1bit](https://huggingface.co/prism-ml/Bonsai-27B-mlx-1bit) for native Apple Silicon inference, including iPhone (\\~11 tok/s on iPhone 17 Pro Max via MLX Swift)\n- **Ternary companion**: the quality-oriented operating point (\\~7.2 GB, 95% of FP16) is also published in GGUF as [Ternary-Bonsai-27B-gguf](https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf)\n\n## Resources\n\n- **[Whitepaper](https://github.com/PrismML-Eng/Bonsai-demo/blob/mai…\n\nSource: https://huggingface.co/prism-ml/Bonsai-27B-gguf","install":{"kind":"model","hfId":"prism-ml/Bonsai-27B-gguf","gated":false,"format":"gguf","files":[{"name":"Bonsai-27B-F16.gguf","size":53808280640,"quant":"F16","sha256":"d4a381a6d07131c34af888607bdbda49fc885c97673a0d22aa3e0f0284bba566"},{"name":"Bonsai-27B-Q1_0.gguf","size":3803452480,"quant":"Q1_0","sha256":"17ef842e47450caeb8eaa3ebfbbab5d2f2278b62b79be107985fb69a2f819aa0"},{"name":"Bonsai-27B-dspark-Q4_1.gguf","size":1787468768,"quant":"Q4_1","sha256":"25e73f9f7ab5d1f7f1336711496dbc12da674e639ec88d579dc8683045befb1b"},{"name":"Bonsai-27B-dspark-bf16.gguf","size":7291885792,"quant":"BF16","sha256":"93bd7f0326bcb39982051dd7742e5a2c9031358f94293a9c91f6aa98dcb6573b"},{"name":"Bonsai-27B-mmproj-BF16.gguf","size":931145760,"quant":"BF16","sha256":"acaf5b55d24ebd38c71fa220dc58c9a36776ec543b17728a4b322fc8d92f1de4"},{"name":"Bonsai-27B-mmproj-Q8_0.gguf","size":629246880,"quant":"Q8_0","sha256":"eb561d41a7bbeb0fcf04883c8af11078ef6cae0a66862a0b68443cfca495269d"}],"totalBytes":68251480320,"suggestedFile":"Bonsai-27B-F16.gguf","requirements":{"ramGb":59,"diskBytes":53808280640,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp","ollama"],"command":"lsh models install hf:prism-ml/Bonsai-27B-gguf"}}