Store › model › LLM
Qwen2.5-Coder-7B-Instruct-GGUF
by Qwen · source Hugging Face · updated 2024-11-12
apache-2.04.4 GB~6 GB RAMsource aliveunlabeled
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 bi…
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:Qwen/Qwen2.5-Coder-7B-Instruct-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct-GGUF
- License: apache-2.0 · text
- Requirements: about 6 GB of RAM, 4.4 GB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp, ollama.
- Tags:
transformersggufcodecodeqwenchatqwenqwen-codertext-generationenendpoints_compatibleconversational
Numbers
- 317,229 downloads on Hugging Face
- 505 likes
- license apache-2.0
- 4.4 GB for qwen2.5-coder-7b-instruct-q4_k_m.gguf
- 27,665 stars on QwenLM/Qwen3
- 68 open issues and PRs
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:53 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
qwen2.5-coder-7b-instruct-fp16-00001-of-00004.gguf | 3.7 GB | |
qwen2.5-coder-7b-instruct-fp16-00002-of-00004.gguf | 3.6 GB | |
qwen2.5-coder-7b-instruct-fp16-00003-of-00004.gguf | 3.6 GB | |
qwen2.5-coder-7b-instruct-fp16-00004-of-00004.gguf | 3.3 GB | |
qwen2.5-coder-7b-instruct-fp16.gguf | 14 GB | |
qwen2.5-coder-7b-instruct-q2_k.gguf | Q2_K | 2.8 GB |
qwen2.5-coder-7b-instruct-q3_k_m.gguf | Q3_K_M | 3.5 GB |
qwen2.5-coder-7b-instruct-q4_0-00001-of-00002.gguf | Q4_0 | 3.7 GB |
qwen2.5-coder-7b-instruct-q4_0-00002-of-00002.gguf | Q4_0 | 427 MB |
qwen2.5-coder-7b-instruct-q4_0.gguf | Q4_0 | 4.1 GB |
qwen2.5-coder-7b-instruct-q4_k_m-00001-of-00002.gguf | Q4_K_M | 3.7 GB |
qwen2.5-coder-7b-instruct-q4_k_m-00002-of-00002.gguf | Q4_K_M | 658 MB |
qwen2.5-coder-7b-instruct-q4_k_m.gguf | Q4_K_M | 4.4 GB |
qwen2.5-coder-7b-instruct-q5_0-00001-of-00002.gguf | Q5_0 | 3.7 GB |
qwen2.5-coder-7b-instruct-q5_0-00002-of-00002.gguf | Q5_0 | 1.2 GB |
qwen2.5-coder-7b-instruct-q5_k_m-00001-of-00002.gguf | Q5_K_M | 3.7 GB |
qwen2.5-coder-7b-instruct-q5_k_m-00002-of-00002.gguf | Q5_K_M | 1.4 GB |
qwen2.5-coder-7b-instruct-q5_k_m.gguf | Q5_K_M | 5.1 GB |
qwen2.5-coder-7b-instruct-q6_k-00001-of-00002.gguf | Q6_K | 3.7 GB |
qwen2.5-coder-7b-instruct-q6_k-00002-of-00002.gguf | Q6_K | 2.1 GB |
qwen2.5-coder-7b-instruct-q6_k.gguf | Q6_K | 5.8 GB |
qwen2.5-coder-7b-instruct-q8_0-00001-of-00003.gguf | Q8_0 | 3.7 GB |
qwen2.5-coder-7b-instruct-q8_0-00002-of-00003.gguf | Q8_0 | 3.7 GB |
qwen2.5-coder-7b-instruct-q8_0-00003-of-00003.gguf | Q8_0 | 167 MB |
qwen2.5-coder-7b-instruct-q8_0.gguf | Q8_0 | 7.5 GB |
From the source README
Qwen2.5-Coder-7B-Instruct-GGUF
Introduction
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings the following improvements upon CodeQwen1.5:
- Significantly improvements in code generation, code reasoning and code fixing. Base on the strong Qwen2.5, we scale up the training tokens into 5.5 trillion including source code, text-code grounding, Synthetic data, etc. Qwen2.5-Coder-32B has become the current state-of-the-art open-source codeLLM, with its coding abilities matching those of GPT-4o.
- A more comprehensive foundation for real-world applications such as Code Agents. Not only enhancing coding capabilities but also maintaining its strengths in mathematics and general competencies.
- Long-context Support up to 128K tokens.
This repo contains the instruction-tuned 7B Qwen2.5-Coder model in the GGUF Format, which has the following features:
- Type: Causal Language Models
- Training Stage: Pretraining & Post-training
- Architecture: transformers with RoPE, SwiGLU, RMSNorm, and Attention QKV bias
- Number of Parameters: 7.61B
- Number of Paramaters (Non-Embedding): 6.53B
- Number of Layers: 28
- Number of Attention Heads (GQA): 28 for Q and 4 for KV
- Context Length: Full 32,768 tokens
- Note: Currently, only vLLM supports YARN for length extrapolating. If you want to process sequences up to 131,072 tokens, please refer to non-GGUF models.
- Quantization: q2_K, q3_K_M, q4_0, q4_K_M, q5_0, q5_K_M, q6_K, q8_0
For more details, please refer to our blog, GitHub, Documentation, [Arxiv](https://arxiv.org/abs…
Source: https://huggingface.co/Qwen/Qwen2.5-Coder-7B-Instruct-GGUF
Card id model:hf:Qwen/Qwen2.5-Coder-7B-Instruct-GGUF · collected 2026-10-02 20:53 UTC · JSON