Store › model › Embeddings
nomic-embed-text-v2-moe-GGUF
by nomic-ai · source Hugging Face · updated 2025-05-15
apache-2.0328 MB~1 GB RAMsource aliveunlabeled
Llama.cpp Quantizations of nomic-embed-text-v2-moe: Multilingual Mixture of Experts Text Embeddings
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:nomic-ai/nomic-embed-text-v2-moe-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/nomic-ai/nomic-embed-text-v2-moe-GGUF
- License: apache-2.0
- Requirements: about 1 GB of RAM, 328 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp.
- Tags:
ggufsentence-similarityfeature-extractionenesfrdeitptplnltrjaviruid
Numbers
- 41,740 downloads on Hugging Face
- 80 likes
- license apache-2.0
- 0.3 GB for nomic-embed-text-v2-moe.Q4_K_M.gguf
- 4,402,710 npm downloads a week for @huggingface/transformers
- latest @huggingface/transformers@4.3.0
Numbers as of 2026-10-02 20:53 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
nomic-embed-text-v2-moe.Q2_K.gguf | Q2_K | 261 MB |
nomic-embed-text-v2-moe.Q3_K_L.gguf | Q3_K_L | 307 MB |
nomic-embed-text-v2-moe.Q3_K_M.gguf | Q3_K_M | 294 MB |
nomic-embed-text-v2-moe.Q3_K_S.gguf | Q3_K_S | 275 MB |
nomic-embed-text-v2-moe.Q4_0.gguf | Q4_0 | 309 MB |
nomic-embed-text-v2-moe.Q4_1.gguf | Q4_1 | 326 MB |
nomic-embed-text-v2-moe.Q4_K_M.gguf | Q4_K_M | 328 MB |
nomic-embed-text-v2-moe.Q4_K_S.gguf | Q4_K_S | 310 MB |
nomic-embed-text-v2-moe.Q5_K_M.gguf | Q5_K_M | 354 MB |
nomic-embed-text-v2-moe.Q5_K_S.gguf | Q5_K_S | 343 MB |
nomic-embed-text-v2-moe.Q6_K.gguf | Q6_K | 379 MB |
nomic-embed-text-v2-moe.Q8_0.gguf | Q8_0 | 488 MB |
nomic-embed-text-v2-moe.bf16.gguf | BF16 | 913 MB |
nomic-embed-text-v2-moe.f16.gguf | F16 | 913 MB |
nomic-embed-text-v2-moe.f32.gguf | F32 | 1.8 GB |
From the source README
Llama.cpp Quantizations of nomic-embed-text-v2-moe: Multilingual Mixture of Experts Text Embeddings
Blog | Technical Report | AWS SageMaker | Atlas Embedding and Unstructured Data Analytics Platform
This model was presented in the paper Training Sparse Mixture Of Experts Text Embedding Models.
Using llama.cpp commit e3a9421b7 for quantization.
Original model: nomic-embed-text-v2-moe
Usage
This model can be used with the llama.cpp server and other software that supports llama.cpp embedding models.
Embedding text with `nomic-embed-text` requires task instruction prefixes at the beginning of each string.
For example, the code below shows how to use the `search_query` prefix to embed user questions, e.g. in a RAG application.
Start a llama.cpp server:
```
llama-server -m nomic-embed-text-v2-moe.bf16.gguf --embeddings
```
And run this code:
```python
import requests
def dot(va, vb):
return sum(a * b for a, b in zip(va, vb))
Card id model:hf:nomic-ai/nomic-embed-text-v2-moe-GGUF · collected 2026-10-02 20:53 UTC · JSON