Store βΊ model βΊ LLM
gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF
by yuxinlu1 Β· source Hugging Face Β· updated 2026-06-19
apache-2.06.9 GB~9 GB RAMsource aliveunlabeled
π»π€ Gemma4-12B v2 β Coding + Agentic Edition β¨ ### π£ Tiny footprint, big brain β a local coding & tool-using agent for everyone
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF
- License: apache-2.0
- Requirements: about 9 GB of RAM, 6.9 GB on disk (estimate: suggested file size Γ 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp, ollama.
- Tags:
ggufgemma4codingagenticterminaltool-usereasoningthinkingllama.cpplocal-llmtext-generationendpoints_compatibleconversational
Numbers
- 710,251 downloads on Hugging Face
- 1,632 likes
- license apache-2.0
- 6.9 GB for gemma4-v2-Q4_K_M.gguf
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:52 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
MTP/gemma-4-12B-it-MTP-BF16.gguf | BF16 | 822 MB |
MTP/gemma-4-12B-it-MTP-F16.gguf | F16 | 822 MB |
MTP/gemma-4-12B-it-MTP-Q8_0.gguf | Q8_0 | 444 MB |
gemma4-v2-Q3_K_M.gguf | Q3_K_M | 5.7 GB |
gemma4-v2-Q4_K_M.gguf | Q4_K_M | 6.9 GB |
gemma4-v2-Q6_K.gguf | Q6_K | 9.1 GB |
gemma4-v2-Q8_0.gguf | Q8_0 | 12 GB |
From the source README
# π»π€ Gemma4-12B v2 β Coding + Agentic Edition β¨
### π£ Tiny footprint, big brain β a local coding & tool-using agent for *everyone*
> No matter your GPU. No matter your RAM. With ~4.5 GB of VRAM *or* unified memory free, you can run your own
> private, offline coding agent right now. π v2 is the big agentic upgrade β it reads, reasons, *uses tools*,
> and works through multi-step technical tasks before it acts. π§ π οΈ All local, all yours, no API, no cloud.
π The headline β it works as an agent (tau2-bench)
v2 is built for coding + agentic work β writing code, running commands, using tools, debugging, multi-step
technical tasks. The clearest signal is tau2-bench `telecom`, an agentic tool-use benchmark whose
*diagnose β fix β verify* loop mirrors real terminal/debugging work:
| tau2-bench telecom Β· 20 tasks Β· local, same harness, all Q8_0 | score |
|---|---|
| official `gemma-4-12B-it` (base) | ~15% |
| π’ Gemma4-12B v2 (this model) | ~55% |
β Roughly 3.5Γ higher than the base model on technical-agentic tasks. π― Want the full story β *why* telecom,
*how* the two models fail differently, the honest caveats, and the trade-offs (including general knowledge)?
It's all broken down further below. π
π Announcements
π Hitting a problem? Please check my pinned discussion first. ~99% of issues are a client/sampler config, not
the weights β and they have a quick fix there. For example: garbled or repeating `0000β¦` output almost always
means no repetition penalty (set `rep_pen 1.1`, `temp 1.0`); and leaked `` / `` tokens mean
your front-end isn't parsing Gemma 4's native tool format (use llama.cpp `--jinja`). If your question isn't covered,
don't hesitate to open a discussion β I read them and reply as fast as I can. π¬
π¦ No Q2_K this release. I finished a Q2β¦
Source: https://huggingface.co/yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF
Card id model:hf:yuxinlu1/gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF Β· collected 2026-10-02 20:52 UTC Β· JSON