{"v":1,"id":"model:hf:unsloth/embeddinggemma-2-GGUF","slug":"model-unsloth-embeddinggemma-2-gguf","kind":"model","category":"embedding","title":"embeddinggemma-2-GGUF","summary":"Read our How to Run EmbeddingGemma 2 Guide!","source":{"provider":"hf","ref":"unsloth/embeddinggemma-2-GGUF","url":"https://huggingface.co/unsloth/embeddinggemma-2-GGUF","rev":"031f0d4b35536f69ab3509d4893c923264fcf253","fetchedAt":"2026-10-09T19:00:58.028Z","etag":"W/\"124c-32G1j6hYvzjx+tFuj/kLSx10rHc\""},"author":{"name":"unsloth","url":"https://huggingface.co/unsloth"},"license":{"spdx":"apache-2.0","raw":"apache-2.0","open":true},"metrics":{"downloads":41582,"downloadsWeek":297971,"likes":212,"stars":5776,"openIssues":330,"lastRelease":{"tag":"v4.0.1","at":"2026-05-20T15:37:55Z"},"pushedAt":"2026-10-06T09:16:03Z","takenAt":"2026-10-09T19:00:58.028Z"},"tags":["gguf","embedding","feature-extraction","multimodal-embedding","multimodal","vision","audio","video","image-feature-extraction","audio-feature-extraction","video-feature-extraction","sentence-similarity","unsloth","multilingual","endpoints_compatible","conversational"],"pipeline":"feature-extraction","links":{"github":"google-deepmind/gemma","npm":"node-llama-cpp"},"updatedAt":"2026-10-07T07:09:29.000Z","collectedAt":"2026-10-09T19:00:58.028Z","review":{"numbers":["41,582 downloads on Hugging Face","212 likes","license apache-2.0","0.3 GB for embeddinggemma-2-Q8_0.gguf","5,776 stars on google-deepmind/gemma","330 open issues and PRs","last release v4.0.1 on 2026-05-20","297,971 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-09T19:00:58.028Z","http":200},"description":"# Read our How to [Run EmbeddingGemma 2 Guide!](https://unsloth.ai/docs/models/embeddinggemma-2)\n\n    Unsloth Dynamic 3.0 achieves superior accuracy & outperforms other leading quants.\n\n    Hugging Face |\n    GitHub |\n    Launch Blog |\n    Documentation |\n\n    License: Apache 2.0 | Authors: Google DeepMind\n\n**EmbeddingGemma 2** is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.\n\nDesigned to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.\n\nEmbeddingGemma 2 builds upon the architectural and capability advancements of Gemma 4, offering several core features:&nbsp;\n\n* **Native multimodality:** Native multimodality: Unifies 4 modalities (text, images, video, and audio) in a single shared 768-dimensional embedding space.\n* **Multilinguality and code:** EmbeddingGemma 2 understands 100+ languages, and achieves a \\~14% improvement on code tasks relative to its predecessor.&nbsp;\n* **Flexible footprint:** Combines a 270M parameter text backbone (130M transformer \\+ 140M embedder) with selectively loadable vision (170M) and audio (300M) encoders, allowing developers to load only the modalities required for their use case.\n* **Matryoshka Representation Learning (MRL):** Native support for truncated embeddings across 128d, 256d, 512d, and 768d, enabling up to a **6x reduction** in vector storage costs with minimal impact on quality.\n* **Context length:** 8K token context window, capable of processing minutes of audio or video.\n* **Task-steered representatio…\n\nSource: https://huggingface.co/unsloth/embeddinggemma-2-GGUF","install":{"kind":"model","hfId":"unsloth/embeddinggemma-2-GGUF","gated":false,"format":"gguf","files":[{"name":"embeddinggemma-2-BF16.gguf","size":557950240,"quant":"BF16","sha256":"f315cbbb30dd487e44d501c8902abe88808755e43753a96beed1964f0a48aa4f"},{"name":"embeddinggemma-2-F16.gguf","size":557950240,"quant":"F16","sha256":"0d4350867eddbb94af485501a3d07e9374c47779d88f8a39cd1abd6a8e2d0901"},{"name":"embeddinggemma-2-Q8_0.gguf","size":309855520,"quant":"Q8_0","sha256":"6f1bd4ac6c5df7444f9cca7ca36cafe6cfa34cd6f49fefb1e0b4be8143aed8bc"},{"name":"embeddinggemma-2-UD-Q4_K_XL.gguf","size":175673856,"quant":"Q4_K_XL","sha256":"ea905fd08e8061db77a0031cbb5096d459d7cfd2bdd0ad2ba66e483f16719493"},{"name":"embeddinggemma-2-UD-Q5_K_XL.gguf","size":210055680,"quant":"Q5_K_XL","sha256":"a9d7a3b7eb421113f05e964e10882676f10fff83811b7876d5bb983db7c7ac41"},{"name":"embeddinggemma-2-UD-Q6_K_XL.gguf","size":248812032,"quant":"Q6_K_XL","sha256":"dc84f042bc3ffbe6122b3b38b2fb299d66f5f37b276c9d458db85d4a4f1e6ba3"},{"name":"mmproj-BF16.gguf","size":982074880,"quant":"BF16","sha256":"995aaa56e88b9b631f651861d659b728a625ccc02a238be37cff56cf11dc0032"},{"name":"mmproj-F16.gguf","size":980895232,"quant":"F16","sha256":"74bf65c860d292a1eebafd7ca85b44532a2528f03017f9204fe1695916d2c66f"},{"name":"mmproj-Q8_0.gguf","size":554821120,"quant":"Q8_0","sha256":"90e7b0238009e2954f856f2081dcf7f35af026b64c765e98f4777053e1754460"}],"totalBytes":4578088800,"suggestedFile":"embeddinggemma-2-Q8_0.gguf","requirements":{"ramGb":1,"diskBytes":309855520,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp"],"command":"lsh models install hf:unsloth/embeddinggemma-2-GGUF"}}