Open models to add in one click
Collected from Hugging Face, GitHub and npm: open LLMs in GGUF, speech models like Whisper, embedding models. Every card links to its source, repeats the license and the numbers, and adds a short summary. One button adds it to LogiShell; without the app it leads to the install and brings you back.
-
Qwen3-Coder-30B-A3B-Instruct-GGUF
model · LLM · by unsloth
See our collection for all versions of Qwen3 including GGUF, 4-bit & 16-bit formats.
apache-2.017 GB~21 GB RAMsource alive
8.6M downloads · 1.1K likes · numbers as of 2026-10-02
-
Ornith-1.5-9B-GGUF
model · LLM · by ornith-ai
Chirp Chirp! 🐦 We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement.
mit5.4 GB~7 GB RAMsource alive
5.1M downloads · 479 likes · numbers as of 2026-10-02
-
Ternary-Bonsai-2-27B-gguf
model · LLM · by prism-ml
Prism ML Website Whitepaper Demo & Examples Discord
apache-2.050 GB~59 GB RAMsource alive
3.9M downloads · 2.4K likes · numbers as of 2026-10-02
-
Ornith-1.5-35B-A3B-GGUF
model · LLM · by ornith-ai
Chirp Chirp! 🐦 We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement.
mit20 GB~24 GB RAMsource alive
3.6M downloads · 473 likes · numbers as of 2026-10-02
-
Ornith-1.0-9B-GGUF
model · LLM · by ornith-ai
Aloha! 🌺 Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding.
mit5.2 GB~7 GB RAMsource alive
2.6M downloads · 673 likes · numbers as of 2026-10-02
-
JiRackUltra_1b
model · LLM · by CMSManhattan
JiRack Ultra 1B (CPU) A fast and efficient ~1.5B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new Routing, Tool call, and Ro…
mit1.0 GB~2 GB RAMsource alive
2.5M downloads · 0 likes · numbers as of 2026-10-02
-
mxbai-embed-large-v1
model · Embeddings · by mixedbread-ai
embedding model by mixedbread-ai, gguf files on Hugging Face.
apache-2.0639 MB~2 GB RAMsource alive
2.0M downloads · 826 likes · numbers as of 2026-10-02
-
nemotron-3.5-asr-streaming-0.6b-gguf
model · Speech · by handy-computer
nemotron-3.5-asr-streaming-0.6b: transcribe.cpp GGUF
other473 MB~2 GB RAMsource alive
1.8M downloads · 12 likes · numbers as of 2026-10-02
-
Ornith-1.0-35B-GGUF
model · LLM · by ornith-ai
Aloha! 🌺 Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding.
mit20 GB~24 GB RAMsource alive
1.7M downloads · 1.1K likes · numbers as of 2026-10-02
-
parakeet-unified-en-0.6b-gguf
model · Speech · by handy-computer
parakeet-unified-en-0.6b: transcribe.cpp GGUF
cc-by-4.0455 MB~2 GB RAMsource alive
1.6M downloads · 9 likes · numbers as of 2026-10-02
-
deepseek-v4-gguf
model · LLM · by antirez
This quants are specific for the DS4 inference engine. They may work with other inference engines or not (they should, but not the MTP model which requires a specific loader).
mit3.5 GB~5 GB RAMsource alive
1.5M downloads · 477 likes · numbers as of 2026-10-02
-
Ornith-1.5-397B-GGUF
model · LLM · by ornith-ai
Chirp Chirp! 🐦 We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement.
mit228 GB~263 GB RAMsource alive
1.3M downloads · 40 likes · numbers as of 2026-10-02
-
nemotron-3.5-asr-streaming-0.6b
model · Speech · by nvidia
This model is the multilingual extension of nvidia/nemotron-speech-streaming-en-0.6b, adding language-ID prompt conditioning to support transcription across 40 language-locales from a single model.
other708 MB~2 GB RAMsource alive
1.2M downloads · 1.2K likes · numbers as of 2026-10-02
-
LFM2.5-2.6B-GGUF
model · LLM · by LiquidAI
LFM2.5 is a new family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with extended pre-training and reinforcement learning.
other1.6 GB~3 GB RAMsource alive
1.1M downloads · 360 likes · numbers as of 2026-10-02
-
GLM-5.3-Flash-GGUF
model · LLM · by unsloth
Unsloth Dynamic 3.0 achieves superior accuracy & outperforms other leading quants.
mit9 MB~1 GB RAMsource alive
1.0M downloads · 443 likes · numbers as of 2026-10-02
-
cohere-transcribe-03-2026-gguf
model · Speech · by handy-computer
cohere-transcribe-03-2026: transcribe.cpp GGUF
apache-2.01.5 GB~3 GB RAMsource alive
978K downloads · 3 likes · numbers as of 2026-10-02
-
JiRackUltra_14b
model · LLM · by CMSManhattan
JiRack Ultra 14B (CPU) A fast and efficient 14B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new Routing, Media, Vision, Sou…
mit8.4 GB~11 GB RAMsource alive
964K downloads · 2 likes · numbers as of 2026-10-02
-
Qwen3.8-27B-NVFP4-MTP-GGUF
model · LLM · by esatapedico
A family of nine GGUF files of Qwen3.8-27B (the native vision-language 27B dense model, Gated DeltaNet + Gated Attention hybrid layout, 262,144-token native context, MTP speculative head), converted…
apache-2.014 GB~17 GB RAMsource alive
752K downloads · 115 likes · numbers as of 2026-10-02
-
gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF
model · LLM · by yuxinlu1
💻🤖 Gemma4-12B v2 — Coding + Agentic Edition ✨ ### 🐣 Tiny footprint, big brain — a local coding & tool-using agent for everyone
apache-2.06.9 GB~9 GB RAMsource alive
710K downloads · 1.6K likes · numbers as of 2026-10-02
-
Qwen3.8-9B-Distill-GGUF
model · LLM · by empero-ai
GGUF quantizations of empero-ai/Qwen3.8-9B — a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-9B architecture — for llama.cpp, Ollama, LM Studio, Jan, KoboldCpp, and other stock GG…
apache-2.05.4 GB~7 GB RAMsource alive
676K downloads · 286 likes · numbers as of 2026-10-02
-
POCKET-35B-GGUF
model · LLM · by FINAL-Bench
🆕 POCKET-Qwen3.8-Flash-Next — a 180B model running on a laptop with 8 GB VRAM + 32 GB RAM · 4.17 tok/s measured. >
apache-2.020 GB~24 GB RAMsource alive
669K downloads · 81 likes · numbers as of 2026-10-02
-
Qwen3.8-4B-Distill-GGUF
model · LLM · by empero-ai
GGUF quantizations of empero-ai/Qwen3.8-4B — a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-4B architecture — for llama.cpp, Ollama, LM Studio, Jan, KoboldCpp, and other stock GG…
apache-2.02.6 GB~4 GB RAMsource alive
666K downloads · 161 likes · numbers as of 2026-10-02
-
Qwen3.8-Flash-Next-GGUF
model · LLM · by AtomicChat
Built from Qwen's original weights with our own importance matrix. The calibration corpora behind our builds are public.
other553 MB~2 GB RAMsource alive
634K downloads · 170 likes · numbers as of 2026-10-02
-
Ternary-Bonsai-27B-gguf
model · LLM · by prism-ml
Prism ML Website Whitepaper Demo & Examples Discord
apache-2.050 GB~59 GB RAMsource alive
628K downloads · 1.4K likes · numbers as of 2026-10-02
136 cards, catalog 27 min oldNext