Store › model › Embeddings
EmbeddingGemma-2-GGUF
by ngquocvinh · source Hugging Face · updated 2026-10-08
apache-2.0173 MB~1 GB RAMsource aliveunlabeled
Community GGUF quantizations of google/embeddinggemma-2.
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:ngquocvinh/EmbeddingGemma-2-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/ngquocvinh/EmbeddingGemma-2-GGUF
- License: apache-2.0
- Requirements: about 1 GB of RAM, 173 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp.
- Tags:
llama.cppggufembeddingmultimodalmultilingualquantizedsentence-transformersfeature-extractionendpoints_compatibleimatrixconversational
Numbers
- 4,365 downloads on Hugging Face
- 1 likes
- license apache-2.0
- 0.2 GB for EmbeddingGemma-2-Q4_K_M.gguf
- 5,776 stars on google-deepmind/gemma
- 330 open issues and PRs
- last release v4.0.1 on 2026-05-20
- 297,971 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-09 19:01 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
EmbeddingGemma-2-IQ1_M.gguf | IQ1_M | 98 MB |
EmbeddingGemma-2-IQ1_S.gguf | IQ1_S | 96 MB |
EmbeddingGemma-2-IQ2_M.gguf | IQ2_M | 125 MB |
EmbeddingGemma-2-IQ2_S.gguf | IQ2_S | 122 MB |
EmbeddingGemma-2-IQ2_XS.gguf | IQ2_XS | 106 MB |
EmbeddingGemma-2-IQ2_XXS.gguf | IQ2_XXS | 102 MB |
EmbeddingGemma-2-IQ3_M.gguf | IQ3_M | 139 MB |
EmbeddingGemma-2-IQ3_S.gguf | IQ3_S | 136 MB |
EmbeddingGemma-2-IQ3_XS.gguf | IQ3_XS | 133 MB |
EmbeddingGemma-2-IQ3_XXS.gguf | IQ3_XXS | 129 MB |
EmbeddingGemma-2-IQ4_NL.gguf | IQ4_NL | 169 MB |
EmbeddingGemma-2-IQ4_XS.gguf | IQ4_XS | 162 MB |
EmbeddingGemma-2-Q1_0.gguf | Q1_0 | 63 MB |
EmbeddingGemma-2-Q2_K.gguf | Q2_K | 115 MB |
EmbeddingGemma-2-Q2_K_S.gguf | Q2_K_S | 111 MB |
EmbeddingGemma-2-Q3_K_L.gguf | Q3_K_L | 147 MB |
EmbeddingGemma-2-Q3_K_M.gguf | Q3_K_M | 142 MB |
EmbeddingGemma-2-Q3_K_S.gguf | Q3_K_S | 136 MB |
EmbeddingGemma-2-Q4_K_M.gguf | Q4_K_M | 173 MB |
EmbeddingGemma-2-Q4_K_S.gguf | Q4_K_S | 170 MB |
EmbeddingGemma-2-Q5_K_M.gguf | Q5_K_M | 203 MB |
EmbeddingGemma-2-Q5_K_S.gguf | Q5_K_S | 201 MB |
EmbeddingGemma-2-Q6_K.gguf | Q6_K | 234 MB |
EmbeddingGemma-2-Q8_0.gguf | Q8_0 | 296 MB |
EmbeddingGemma-2-TQ1_0.gguf | TQ1_0 | 126 MB |
EmbeddingGemma-2-TQ2_0.gguf | TQ2_0 | 132 MB |
mmproj-EmbeddingGemma-2-BF16.gguf | BF16 | 937 MB |
From the source README
EmbeddingGemma 2 GGUF
Community GGUF quantizations of google/embeddinggemma-2.
☕ If this GGUF made your day easier, a coffee would make mine.
Send a coffee ☕
I build and test these releases myself. Your coffee helps keep me going.
Thank you for supporting this work.
About EmbeddingGemma 2
EmbeddingGemma 2 is a multilingual, multimodal embedding model from Google DeepMind. It maps text and code, images, video, and audio into a shared 768-dimensional vector space and has an 8,192-token shared context window. The upstream checkpoint has a 270M-parameter text path plus optional vision and audio encoders. This release contains a quantized text backbone and a shared BF16 vision/audio projector. Text, image, and audio inputs returned normalized 768-dimensional vectors in local `llama-server` CPU smoke checks. See the official model card for supported task prefixes, modalities, and input guidance.
Embedding evaluation
Every value below comes from local measurements of the locked BF16 GGUF and quantized files; the numbers are not copied from the upstream model card. The fixed task evaluation uses the 1,379-pair test split of MTEB STSBenchmark STS, dataset revision `96943a16ea6a35129e253c659081cb59daf81b30`. It covers 2,552 unique sentences with the `task: sentence similarity | query:` prefix, an 8,192-token context, and the CPU `llama.cpp` runtime at commit `9c2e0e491a822adae1f0b1c831adb4160057d24f`. Spearman and Pearson measure correlation between cosine similarity and human scores. Mean and 5th-percentile cosine measure vector agreement with the BF16 GGUF reference. Higher values indicate stronger task correlation or closer vector agreement; smaller files use less disk. This is a task-specif…
Source: https://huggingface.co/ngquocvinh/EmbeddingGemma-2-GGUF
Card id model:hf:ngquocvinh/EmbeddingGemma-2-GGUF · collected 2026-10-09 19:01 UTC · JSON