Store › model › Embeddings
Qwen3-Embedding-0.6B-GGUF
by PeterAM4 · source Hugging Face · updated 2026-02-07
apache-2.0378 MB~1 GB RAMsource aliveunlabeled
All-in-one GGUF quantizations of Qwen/Qwen3-Embedding-0.6B, from 8-bit down to 1-bit, with importance-matrix calibration optimized for financial and technical text retrieval.
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:PeterAM4/Qwen3-Embedding-0.6B-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/PeterAM4/Qwen3-Embedding-0.6B-GGUF
- License: apache-2.0
- Requirements: about 1 GB of RAM, 378 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp.
- Tags:
ggufquantizedembeddingsentence-transformersQwen3imatrixllama-cppsentence-similarityenzhjakofrdeespt
Numbers
- 2,591 downloads on Hugging Face
- 3 likes
- license apache-2.0
- 0.4 GB for Qwen3-Embedding-0.6B-Q4_K_M-imat.gguf
- 27,665 stars on QwenLM/Qwen3
- 68 open issues and PRs
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:53 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
Qwen3-Embedding-0.6B-BF16.gguf | BF16 | 1.1 GB |
Qwen3-Embedding-0.6B-IQ1_M-imat.gguf | IQ1_M | 206 MB |
Qwen3-Embedding-0.6B-IQ1_S-imat.gguf | IQ1_S | 198 MB |
Qwen3-Embedding-0.6B-IQ2_M-imat.gguf | IQ2_M | 252 MB |
Qwen3-Embedding-0.6B-IQ2_S-imat.gguf | IQ2_S | 242 MB |
Qwen3-Embedding-0.6B-IQ2_XS-imat.gguf | IQ2_XS | 231 MB |
Qwen3-Embedding-0.6B-IQ2_XXS-imat.gguf | IQ2_XXS | 219 MB |
Qwen3-Embedding-0.6B-IQ3_M-imat.gguf | IQ3_M | 320 MB |
Qwen3-Embedding-0.6B-IQ3_S-imat.gguf | IQ3_S | 308 MB |
Qwen3-Embedding-0.6B-IQ3_XS-imat.gguf | IQ3_XS | 298 MB |
Qwen3-Embedding-0.6B-IQ3_XXS-imat.gguf | IQ3_XXS | 266 MB |
Qwen3-Embedding-0.6B-IQ4_NL-imat.gguf | IQ4_NL | 364 MB |
Qwen3-Embedding-0.6B-IQ4_XS-imat.gguf | IQ4_XS | 351 MB |
Qwen3-Embedding-0.6B-Q2_K-imat.gguf | Q2_K | 282 MB |
Qwen3-Embedding-0.6B-Q2_K_S-imat.gguf | Q2_K_S | 267 MB |
Qwen3-Embedding-0.6B-Q3_K_L-imat.gguf | Q3_K_L | 351 MB |
Qwen3-Embedding-0.6B-Q3_K_M-imat.gguf | Q3_K_M | 331 MB |
Qwen3-Embedding-0.6B-Q3_K_S-imat.gguf | Q3_K_S | 308 MB |
Qwen3-Embedding-0.6B-Q4_0-imat.gguf | Q4_0 | 364 MB |
Qwen3-Embedding-0.6B-Q4_1-imat.gguf | Q4_1 | 390 MB |
Qwen3-Embedding-0.6B-Q4_K_M-imat.gguf | Q4_K_M | 378 MB |
Qwen3-Embedding-0.6B-Q4_K_S-imat.gguf | Q4_K_S | 365 MB |
Qwen3-Embedding-0.6B-Q5_0.gguf | Q5_0 | 416 MB |
Qwen3-Embedding-0.6B-Q5_1.gguf | Q5_1 | 442 MB |
Qwen3-Embedding-0.6B-Q5_K_M.gguf | Q5_K_M | 424 MB |
Qwen3-Embedding-0.6B-Q5_K_S.gguf | Q5_K_S | 416 MB |
Qwen3-Embedding-0.6B-Q6_K.gguf | Q6_K | 472 MB |
Qwen3-Embedding-0.6B-Q8_0.gguf | Q8_0 | 610 MB |
Qwen3-Embedding-0.6B-TQ1_0-imat.gguf | TQ1_0 | 216 MB |
Qwen3-Embedding-0.6B-TQ2_0-imat.gguf | TQ2_0 | 236 MB |
From the source README
Qwen3-Embedding-0.6B -- GGUF
All-in-one GGUF quantizations of Qwen/Qwen3-Embedding-0.6B, from 8-bit down to 1-bit, with importance-matrix calibration optimized for financial and technical text retrieval.
Qwen3-Embedding-0.6B is a compact, multilingual embedding model well suited for RAG pipelines, semantic search, and document retrieval. These quantizations make it practical to run on edge devices, laptops, and resource-constrained servers -- particularly for financial NLP workloads where low latency and small memory footprint matter.
The importance matrix was calibrated on a mixed corpus weighted toward financial data (financial Q&A from FiQA, SEC 10-K filings from FinanceBench, financial sentiment from Twitter, RAG pairs) alongside math reasoning and general text, so the quantized models preserve the weights most relevant to financial domain embeddings.
| Property | Value |
|----------|-------|
| Base model | Qwen/Qwen3-Embedding-0.6B |
| Parameters | 595,776,512 |
| Max context | 32,768 tokens |
| Pooling | Last token |
| Embedding dim | 1024 |
| License | Apache 2.0 |
| Quantized with | llama.cpp |
Why quantize?
Despite their size, neural networks are remarkably sparse in information density. Most of the 16 bits allocated per weight during training exist to make gradient descent work -- not to store knowledge. Current estimates put the actual information content at roughly 2 bits per parameter. The remaining 14 bits are redundancy.
This explains why aggressive quantization works: compressing from 16-bit to 4-bit (75% reduction) discards almost exclusively noise. Our benchmark data confirms this -- Q3_K_M-imat at 4.66 BPW scores within +0.62 PPL of the full BF16 baseline while being 70% smaller than BF16 (331 MB vs 1.1 GB). For comparis…
Source: https://huggingface.co/PeterAM4/Qwen3-Embedding-0.6B-GGUF
Card id model:hf:PeterAM4/Qwen3-Embedding-0.6B-GGUF · collected 2026-10-02 20:53 UTC · JSON