LogiShell store Open app

Store › model › Embeddings

Qwen3-Embedding-0.6B-GGUF

by PeterAM4 · source Hugging Face · updated 2026-02-07

apache-2.0378 MB~1 GB RAMsource aliveunlabeled

All-in-one GGUF quantizations of Qwen/Qwen3-Embedding-0.6B, from 8-bit down to 1-bit, with importance-matrix calibration optimized for financial and technical text retrieval.

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:PeterAM4/Qwen3-Embedding-0.6B-GGUF and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-02 20:53 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
Qwen3-Embedding-0.6B-BF16.ggufBF161.1 GB
Qwen3-Embedding-0.6B-IQ1_M-imat.ggufIQ1_M206 MB
Qwen3-Embedding-0.6B-IQ1_S-imat.ggufIQ1_S198 MB
Qwen3-Embedding-0.6B-IQ2_M-imat.ggufIQ2_M252 MB
Qwen3-Embedding-0.6B-IQ2_S-imat.ggufIQ2_S242 MB
Qwen3-Embedding-0.6B-IQ2_XS-imat.ggufIQ2_XS231 MB
Qwen3-Embedding-0.6B-IQ2_XXS-imat.ggufIQ2_XXS219 MB
Qwen3-Embedding-0.6B-IQ3_M-imat.ggufIQ3_M320 MB
Qwen3-Embedding-0.6B-IQ3_S-imat.ggufIQ3_S308 MB
Qwen3-Embedding-0.6B-IQ3_XS-imat.ggufIQ3_XS298 MB
Qwen3-Embedding-0.6B-IQ3_XXS-imat.ggufIQ3_XXS266 MB
Qwen3-Embedding-0.6B-IQ4_NL-imat.ggufIQ4_NL364 MB
Qwen3-Embedding-0.6B-IQ4_XS-imat.ggufIQ4_XS351 MB
Qwen3-Embedding-0.6B-Q2_K-imat.ggufQ2_K282 MB
Qwen3-Embedding-0.6B-Q2_K_S-imat.ggufQ2_K_S267 MB
Qwen3-Embedding-0.6B-Q3_K_L-imat.ggufQ3_K_L351 MB
Qwen3-Embedding-0.6B-Q3_K_M-imat.ggufQ3_K_M331 MB
Qwen3-Embedding-0.6B-Q3_K_S-imat.ggufQ3_K_S308 MB
Qwen3-Embedding-0.6B-Q4_0-imat.ggufQ4_0364 MB
Qwen3-Embedding-0.6B-Q4_1-imat.ggufQ4_1390 MB
Qwen3-Embedding-0.6B-Q4_K_M-imat.ggufQ4_K_M378 MB
Qwen3-Embedding-0.6B-Q4_K_S-imat.ggufQ4_K_S365 MB
Qwen3-Embedding-0.6B-Q5_0.ggufQ5_0416 MB
Qwen3-Embedding-0.6B-Q5_1.ggufQ5_1442 MB
Qwen3-Embedding-0.6B-Q5_K_M.ggufQ5_K_M424 MB
Qwen3-Embedding-0.6B-Q5_K_S.ggufQ5_K_S416 MB
Qwen3-Embedding-0.6B-Q6_K.ggufQ6_K472 MB
Qwen3-Embedding-0.6B-Q8_0.ggufQ8_0610 MB
Qwen3-Embedding-0.6B-TQ1_0-imat.ggufTQ1_0216 MB
Qwen3-Embedding-0.6B-TQ2_0-imat.ggufTQ2_0236 MB

From the source README

Qwen3-Embedding-0.6B -- GGUF

All-in-one GGUF quantizations of Qwen/Qwen3-Embedding-0.6B, from 8-bit down to 1-bit, with importance-matrix calibration optimized for financial and technical text retrieval.

Qwen3-Embedding-0.6B is a compact, multilingual embedding model well suited for RAG pipelines, semantic search, and document retrieval. These quantizations make it practical to run on edge devices, laptops, and resource-constrained servers -- particularly for financial NLP workloads where low latency and small memory footprint matter.

The importance matrix was calibrated on a mixed corpus weighted toward financial data (financial Q&A from FiQA, SEC 10-K filings from FinanceBench, financial sentiment from Twitter, RAG pairs) alongside math reasoning and general text, so the quantized models preserve the weights most relevant to financial domain embeddings.

| Property | Value |
|----------|-------|
| Base model | Qwen/Qwen3-Embedding-0.6B |
| Parameters | 595,776,512 |
| Max context | 32,768 tokens |
| Pooling | Last token |
| Embedding dim | 1024 |
| License | Apache 2.0 |
| Quantized with | llama.cpp |

Why quantize?

Despite their size, neural networks are remarkably sparse in information density. Most of the 16 bits allocated per weight during training exist to make gradient descent work -- not to store knowledge. Current estimates put the actual information content at roughly 2 bits per parameter. The remaining 14 bits are redundancy.

This explains why aggressive quantization works: compressing from 16-bit to 4-bit (75% reduction) discards almost exclusively noise. Our benchmark data confirms this -- Q3_K_M-imat at 4.66 BPW scores within +0.62 PPL of the full BF16 baseline while being 70% smaller than BF16 (331 MB vs 1.1 GB). For comparis…

Source: https://huggingface.co/PeterAM4/Qwen3-Embedding-0.6B-GGUF

Card id model:hf:PeterAM4/Qwen3-Embedding-0.6B-GGUF · collected 2026-10-02 20:53 UTC · JSON