Store › model › Embeddings
embeddinggemma-2-GGUF
by unsloth · source Hugging Face · updated 2026-10-07
apache-2.0296 MB~1 GB RAMsource aliveunlabeled
Read our How to Run EmbeddingGemma 2 Guide!
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:unsloth/embeddinggemma-2-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/unsloth/embeddinggemma-2-GGUF
- License: apache-2.0
- Requirements: about 1 GB of RAM, 296 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp.
- Tags:
ggufembeddingfeature-extractionmultimodal-embeddingmultimodalvisionaudiovideoimage-feature-extractionaudio-feature-extractionvideo-feature-extractionsentence-similarityunslothmultilingualendpoints_compatibleconversational
Numbers
- 41,582 downloads on Hugging Face
- 212 likes
- license apache-2.0
- 0.3 GB for embeddinggemma-2-Q8_0.gguf
- 5,776 stars on google-deepmind/gemma
- 330 open issues and PRs
- last release v4.0.1 on 2026-05-20
- 297,971 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-09 19:00 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
embeddinggemma-2-BF16.gguf | BF16 | 532 MB |
embeddinggemma-2-F16.gguf | F16 | 532 MB |
embeddinggemma-2-Q8_0.gguf | Q8_0 | 296 MB |
embeddinggemma-2-UD-Q4_K_XL.gguf | Q4_K_XL | 168 MB |
embeddinggemma-2-UD-Q5_K_XL.gguf | Q5_K_XL | 200 MB |
embeddinggemma-2-UD-Q6_K_XL.gguf | Q6_K_XL | 237 MB |
mmproj-BF16.gguf | BF16 | 937 MB |
mmproj-F16.gguf | F16 | 935 MB |
mmproj-Q8_0.gguf | Q8_0 | 529 MB |
From the source README
Read our How to Run EmbeddingGemma 2 Guide!
Unsloth Dynamic 3.0 achieves superior accuracy & outperforms other leading quants.
Hugging Face |
GitHub |
Launch Blog |
Documentation |
License: Apache 2.0 | Authors: Google DeepMind
EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.
Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.
EmbeddingGemma 2 builds upon the architectural and capability advancements of Gemma 4, offering several core features:
- Native multimodality: Native multimodality: Unifies 4 modalities (text, images, video, and audio) in a single shared 768-dimensional embedding space.
- Multilinguality and code: EmbeddingGemma 2 understands 100+ languages, and achieves a \~14% improvement on code tasks relative to its predecessor.
- Flexible footprint: Combines a 270M parameter text backbone (130M transformer \+ 140M embedder) with selectively loadable vision (170M) and audio (300M) encoders, allowing developers to load only the modalities required for their use case.
- Matryoshka Representation Learning (MRL): Native support for truncated embeddings across 128d, 256d, 512d, and 768d, enabling up to a 6x reduction in vector storage costs with minimal impact on quality.
- Context length: 8K token context window, capable of processing minutes of audio or video.
- **Task-steered representatio…
Source: https://huggingface.co/unsloth/embeddinggemma-2-GGUF
Card id model:hf:unsloth/embeddinggemma-2-GGUF · collected 2026-10-09 19:00 UTC · JSON