Store βΊ model βΊ LLM
gemma-4-12B-coder-fable5-composer2.5-v1-GGUF
by yuxinlu1 Β· source Hugging Face Β· updated 2026-06-19
apache-2.06.9 GB~9 GB RAMsource aliveunlabeled
π» Gemma4-12B-Coder (GGUF) β Composer 2.5 Γ Fable 5 β¨ ### π£ Tiny footprint, big brain β a local coding model for everyone
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF
- License: apache-2.0
- Requirements: about 9 GB of RAM, 6.9 GB on disk (estimate: suggested file size Γ 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp, ollama.
- Tags:
ggufgemma4codingcodereasoningthinkingllama.cpplocal-llmtext-generationendpoints_compatibleconversational
Numbers
- 405,736 downloads on Hugging Face
- 2,919 likes
- license apache-2.0
- 6.9 GB for gemma4-coding-Q4_K_M.gguf
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:53 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
gemma4-coding-Q2_K.gguf | Q2_K | 4.5 GB |
gemma4-coding-Q3_K_M.gguf | Q3_K_M | 5.7 GB |
gemma4-coding-Q4_K_M.gguf | Q4_K_M | 6.9 GB |
gemma4-coding-Q6_K.gguf | Q6_K | 9.1 GB |
gemma4-coding-Q8_0.gguf | Q8_0 | 12 GB |
From the source README
# π» Gemma4-12B-Coder (GGUF) β Composer 2.5 Γ Fable 5 β¨
### π£ Tiny footprint, big brain β a local coding model for *everyone*
> No matter your GPU. No matter your RAM. If you've got ~4.5 GB of VRAM *or* unified memory free,
> you can run your own private, offline coding assistant right now. π
> This is the v1 / code edition β distilled from real chain-of-thought so it *thinks through* a problem
> before writing the solution. π§ π» All local, all yours, no API, no cloud.
### π― What it is
A focused fine-tune of Gemma 4 12B on verifiable Python coding data β every training example's reasoning leads to
code that actually passed its tests. The result reasons in the open (edge cases, complexity, approach) and then
emits a clean, runnable solution. π
π Announcements
ππ₯ IT'S HERE β v2 is OUT NOW! v2 has shipped β the GGUF quants are live and ready to run β
grab v2 here. π
The full `safetensors` master (build / fine-tune on top) goes up tomorrow. v2 is agentic + coding focused β
the piece v1 was missing.
Here's the result that got me most excited. When I saw v2's tau2-bench `telecom` result β an agentic tool-use
benchmark where the model has to *diagnose β fix β verify*, exactly like real terminal/debugging work β I literally got
launched out of my chair (β¦okay, *kidding* π). The jump in actually solving the problem is wild:
| tau2-bench telecom Β· local, same harness, Q8_0 | score |
|---|---|
| official `gemma-4-12B-it` (base) | ~15% |
| π’ v2 (this release) | ~55% |
The base model tends to give up early (hands the problem off to a human); v2 keeps going and works it the way a
much bigger model would. Full benchmark details are in the **[v2 card](https://huggingface.co/yuxinlu1/gemmβ¦
Source: https://huggingface.co/yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF
Card id model:hf:yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF Β· collected 2026-10-02 20:53 UTC Β· JSON