Store › model › LLM
Qwen3-0.6B-GGUF
by Qwen · source Hugging Face · updated 2025-05-09
apache-2.0610 MB~2 GB RAMsource aliveunlabeled
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers grou…
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:Qwen/Qwen3-0.6B-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/Qwen/Qwen3-0.6B-GGUF
- License: apache-2.0 · text
- Requirements: about 2 GB of RAM, 610 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp, ollama.
- Tags:
gguftext-generationendpoints_compatibleconversational
Numbers
- 345,227 downloads on Hugging Face
- 88 likes
- license apache-2.0
- 0.6 GB for Qwen3-0.6B-Q8_0.gguf
- 27,665 stars on QwenLM/Qwen3
- 68 open issues and PRs
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:53 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
Qwen3-0.6B-Q8_0.gguf | Q8_0 | 610 MB |
From the source README
Qwen3-0.6B-GGUF
Qwen3 Highlights
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:
- Uniquely support of seamless switching between thinking mode (for complex logical reasoning, math, and coding) and non-thinking mode (for efficient, general-purpose dialogue) within single model, ensuring optimal performance across various scenarios.
- Significantly enhancement in its reasoning capabilities, surpassing previous QwQ (in thinking mode) and Qwen2.5 instruct models (in non-thinking mode) on mathematics, code generation, and commonsense logical reasoning.
- Superior human preference alignment, excelling in creative writing, role-playing, multi-turn dialogues, and instruction following, to deliver a more natural, engaging, and immersive conversational experience.
- Expertise in agent capabilities, enabling precise integration with external tools in both thinking and unthinking modes and achieving leading performance among open-source models in complex agent-based tasks.
- Support of 100+ languages and dialects with strong capabilities for multilingual instruction following and translation.
Model Overview
Qwen3-0.6B has the following features:
- Type: Causal Language Models
- Training Stage: Pretraining & Post-training
- Number of Parameters: 0.6B
- Number of Paramaters (Non-Embedding): 0.44B
- Number of Layers: 28
- Number of Attention Heads (GQA): 16 for Q and 8 for KV
- Context Length: 32,768.
- Quantization: q8_0
For more details, including benchmark evaluation, hardware requirements, and inference performance, please refer to our [blog](https://qwenlm.github.io/blog/qwen3…
Source: https://huggingface.co/Qwen/Qwen3-0.6B-GGUF
Card id model:hf:Qwen/Qwen3-0.6B-GGUF · collected 2026-10-02 20:53 UTC · JSON