Store › model › LLM
Ternary-Bonsai-27B-gguf
by prism-ml · source Hugging Face · updated 2026-08-31
apache-2.050 GB~59 GB RAMsource aliveunlabeled
Prism ML Website Whitepaper Demo & Examples Discord
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:prism-ml/Ternary-Bonsai-27B-gguf and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf
- License: apache-2.0
- Requirements: about 59 GB of RAM, 50 GB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp, ollama.
- Tags:
llama.cppggufconversationalternary2-bitllama-cppcudametalon-devicehybrid-attentionprismmlbonsaitext-generationeval-resultsendpoints_compatible
Numbers
- 628,497 downloads on Hugging Face
- 1,403 likes
- license apache-2.0
- 50 GB for Ternary-Bonsai-27B-F16.gguf
- 130,156 stars on ggml-org/llama.cpp
- 2,528 open issues and PRs
- last release v0.5.0 on 2026-09-23
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:52 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
Ternary-Bonsai-27B-F16.gguf | F16 | 50 GB |
Ternary-Bonsai-27B-PQ2_0.gguf | Q2_0 | 6.7 GB |
Ternary-Bonsai-27B-Q2_0.gguf | Q2_0 | 6.7 GB |
Ternary-Bonsai-27B-Q2_g64.gguf | Q2_G | 7.1 GB |
Ternary-Bonsai-27B-dspark-Q4_1.gguf | Q4_1 | 1.8 GB |
Ternary-Bonsai-27B-dspark-bf16.gguf | BF16 | 6.8 GB |
Ternary-Bonsai-27B-mmproj-BF16.gguf | BF16 | 888 MB |
Ternary-Bonsai-27B-mmproj-Q8_0.gguf | Q8_0 | 600 MB |
From the source README
Prism ML Website |
Whitepaper |
Demo & Examples |
Discord
Ternary Bonsai 27B — GGUF
Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU)
> \~9.4x smaller than FP16 (ideal) | 95% of FP16 intelligence retained | \~26 tok/s on an Apple M5 Pro laptop
Highlights
- \~7.2 GB deployed footprint (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU
- 95% of FP16 intelligence retained: 80.49 average across 15 thinking-mode benchmarks — a *higher* score than the conventional IQ2_XXS build (72.73) at less than two-thirds of its footprint
- Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within two points of full precision (93.40), coding at 85.96, agentic tool use at 74.01
- End-to-end ternary language weights across embeddings, attention projections, MLP projections, and LM head, at a *true* 1.71 bits per weight — no high-precision escape hatches behind a low-bit label; the vision tower ships in compact 4-bit HQQ
- 262K-token context on-device, kept practical by the Qwen3.6-27B hybrid-attention backbone (\~75% linear attention) and 4-bit KV-cache quantization
- GGUF Q2_0_g128 format with custom 2-bit hybrid-attention kernels for llama.cpp (CUDA, Metal) — packed weights are consumed directly, never expanded back to FP16
- Ships with a DSpark speculative-decoding drafter layer trained against the Bonsai 27B target — a lossless 1.34x decode speedup on the CUDA serving path
- MLX companion: also available as Ternary-Bonsai-27B-mlx-2bit for native Apple Silicon inference
- 1-bit companion: the phone-class operating point (\~3.9 GB) that fits an iPhone 17 Pro Max, published in GGUF as [Bonsa…
Source: https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf
Card id model:hf:prism-ml/Ternary-Bonsai-27B-gguf · collected 2026-10-02 20:52 UTC · JSON