Store › model › LLM
Ternary-Bonsai-8B-gguf
by prism-ml · source Hugging Face · updated 2026-06-10
apache-2.015 GB~19 GB RAMsource aliveunlabeled
Prism ML Website White Paper Demo & Examples Discord
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:prism-ml/Ternary-Bonsai-8B-gguf and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/prism-ml/Ternary-Bonsai-8B-gguf
- License: apache-2.0
- Requirements: about 19 GB of RAM, 15 GB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp, ollama.
- Tags:
ggufternary1.58-bitllama-cppq2_0on-deviceprismmlbonsaitext-generationeval-resultsendpoints_compatibleconversational
Numbers
- 371,478 downloads on Hugging Face
- 164 likes
- license apache-2.0
- 15 GB for Ternary-Bonsai-8B-F16.gguf
- 130,156 stars on ggml-org/llama.cpp
- 2,528 open issues and PRs
- last release v0.5.0 on 2026-09-23
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:53 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
Ternary-Bonsai-8B-F16.gguf | F16 | 15 GB |
Ternary-Bonsai-8B-PQ2_0.gguf | Q2_0 | 2.0 GB |
Ternary-Bonsai-8B-Q2_0.gguf | Q2_0 | 2.0 GB |
Ternary-Bonsai-8B-Q2_0_g64.gguf | Q2_0 | 2.2 GB |
From the source README
Prism ML Website |
White Paper |
Demo & Examples |
Discord
Ternary-Bonsai-8B-gguf
Ternary (1.58-bit) language model in GGUF Q2_0 format for `llama.cpp`
Resources
- White Paper
- Demo repo — examples for serving, benchmarking, and integrating Bonsai
- Discord — community support and updates
- Kernels: Q2_0 is not yet in mainline `llama.cpp`. Use our fork at PrismML-Eng/llama.cpp (`prism` branch, default) which adds Q2_0 support for CPU (NEON/generic) and Metal. Upstream PR coming soon.
Model Overview
| Item | Specification |
| :--------------- | :----------------------------------------------------------------------- |
| Base model | Qwen3-8B |
| Parameters | 8.19B (~6.95B non-embedding) |
| Architecture | GQA (32 query / 8 KV heads), SwiGLU MLP, RoPE, RMSNorm |
| Layers | 36 Transformer decoder blocks |
| Context length | 65,536 tokens |
| Vocab size | 151,936 |
| Weight format | GGUF Q2_0 g128: {-1, 0, +1} with FP16 group-wise scaling |
| Packed Q2_0 size | 2.03 GiB (2.18 GB) |
| Ternary coverage | Embeddings, attention projections, MLP projections, LM head |
| License | Apache 2.0…
Source: https://huggingface.co/prism-ml/Ternary-Bonsai-8B-gguf
Card id model:hf:prism-ml/Ternary-Bonsai-8B-gguf · collected 2026-10-02 20:53 UTC · JSON