LogiShell store Open app

Store › model › LLM

Bonsai-27B-gguf

by prism-ml · source Hugging Face · updated 2026-07-17

apache-2.050 GB~59 GB RAMsource aliveunlabeled

Prism ML Website Whitepaper Demo & Examples Discord

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:prism-ml/Bonsai-27B-gguf and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-02 20:53 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
Bonsai-27B-F16.ggufF1650 GB
Bonsai-27B-Q1_0.ggufQ1_03.5 GB
Bonsai-27B-dspark-Q4_1.ggufQ4_11.7 GB
Bonsai-27B-dspark-bf16.ggufBF166.8 GB
Bonsai-27B-mmproj-BF16.ggufBF16888 MB
Bonsai-27B-mmproj-Q8_0.ggufQ8_0600 MB

From the source README

Prism ML Website  | 
Whitepaper  | 
Demo & Examples  | 
Discord

1-bit Bonsai 27B — GGUF

Full 27B-class reasoning in binary transformer weights, for llama.cpp (CUDA, Metal, CPU)

> \~14.2x smaller than FP16 | \~90% of FP16 intelligence retained | \~44 tok/s on an Apple M5 Pro laptop

Highlights

  • \~3.9 GB deployed footprint (down from \~54 GB FP16) — a 27B model on everyday laptops and single GPUs
  • Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse — 76.11 average across 15 thinking-mode benchmarks (89.5% of FP16), including math at 91.66 and coding at 81.88
  • End-to-end binary language weights across embeddings, attention projections, MLP projections, and LM head, at a *true* 1.125 bits per weight — no high-precision escape hatches behind a low-bit label; the vision tower ships in compact 4-bit HQQ
  • 262K-token context on-device, kept practical by the Qwen3.6-27B hybrid-attention backbone (\~75% linear attention) and 4-bit KV-cache quantization
  • GGUF Q1_0_g128 format with custom 1-bit hybrid-attention kernels for llama.cpp (CUDA, Metal) — packed weights are consumed directly, never expanded back to FP16
  • Ships with a DSpark speculative-decoding drafter layer trained against the Bonsai 27B target — a lossless 1.37x decode speedup on the CUDA serving path
  • MLX companion: also available as Bonsai-27B-mlx-1bit for native Apple Silicon inference, including iPhone (\~11 tok/s on iPhone 17 Pro Max via MLX Swift)
  • Ternary companion: the quality-oriented operating point (\~7.2 GB, 95% of FP16) is also published in GGUF as Ternary-Bonsai-27B-gguf

Resources

  • **[Whitepaper](https://github.com/PrismML-Eng/Bonsai-demo/blob/mai…

Source: https://huggingface.co/prism-ml/Bonsai-27B-gguf

Card id model:hf:prism-ml/Bonsai-27B-gguf · collected 2026-10-02 20:53 UTC · JSON