LogiShell store Open app

Store › model › LLM

Ternary-Bonsai-2-27B-gguf

by prism-ml · source Hugging Face · updated 2026-09-25

apache-2.050 GB~59 GB RAMsource aliveunlabeled

Prism ML Website Whitepaper Demo & Examples Discord

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:prism-ml/Ternary-Bonsai-2-27B-gguf and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-02 20:54 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
Ternary-Bonsai-2-27B-F16.ggufF1650 GB
Ternary-Bonsai-2-27B-PQ2_0.ggufQ2_06.7 GB
Ternary-Bonsai-2-27B-PTQ1_0.ggufTQ1_05.5 GB
Ternary-Bonsai-2-27B-mmproj-BF16.ggufBF16888 MB
Ternary-Bonsai-2-27B-mmproj-Q8_0.ggufQ8_0600 MB

From the source README

Prism ML Website  | 
Whitepaper  | 
Demo & Examples  | 
Discord

Bonsai 2 27B — GGUF

Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU)

> \~9.3x smaller than FP16 (ideal) | 98.2% of FP16 intelligence retained | \~47 tok/s on an Apple M5 Max laptop

Highlights

  • \~5.9 GB language model (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU
  • 98.2% of FP16 intelligence retained: 84.78 average across 14 thinking-mode benchmarks — far above the conventional IQ2_XXS build (72.59) at about 82% of its footprint, and within 0.4 points of UD-Q4_K_XL at three times the footprint
  • Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within half a point of full precision (96.57), coding level with the baseline (89.42), agentic tool calling at 74.92
  • End-to-end ternary language weights across embeddings, attention projections, MLP projections, and LM head, at a *true* 1.72 bits per weight — no high-precision escape hatches behind a low-bit label; the vision tower ships as a separate Q8_0 mmproj pack
  • 262K-token context on-device, kept practical by the Qwen3.8-27B hybrid-attention backbone (\~75% linear attention)
  • Two GGUF packings with custom ternary hybrid-attention kernels for llama.cpp (CUDA, Metal) — PTQ1_0 packs trits densely (1.75 bits/weight, 5.95 GB), PQ2_0 stores each trit in a 2-bit slot (2.13 bits/weight, 7.21 GB); packed weights are consumed directly, never expanded back to FP16
  • MLX companion: also available as Ternary-Bonsai-2-27B-mlx-2bit for native Apple Silicon inference

Resources

  • **[Whitepaper](https://github.com/PrismML-Eng/Bonsai-demo/blob/main/bonsai-2-27b-whitepape…

Source: https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf

Card id model:hf:prism-ml/Ternary-Bonsai-2-27B-gguf · collected 2026-10-02 20:54 UTC · JSON