Store › model › Embeddings
jina-embeddings-v5-text-nano-retrieval-GGUF
by jinaai · source Hugging Face · updated 2026-02-27
cc-by-nc-4.0150 MB~1 GB RAMsource aliveunlabeled
jina-embeddings-v5-text-nano-retrieval-GGUF
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:jinaai/jina-embeddings-v5-text-nano-retrieval-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/jinaai/jina-embeddings-v5-text-nano-retrieval-GGUF
- License: cc-by-nc-4.0 (restricted license: terms at the source apply)
- Requirements: about 1 GB of RAM, 150 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp.
- Tags:
llama.cppggufembeddingeurobertllama-cppjina-embeddings-v5sentence-similaritymultilingualfeature-extraction
Numbers
- 4,575 downloads on Hugging Face
- 4 likes
- license cc-by-nc-4.0
- 0.1 GB for v5-nano-retrieval-Q4_K_M.gguf
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 21:00 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
v5-nano-retrieval-F16.gguf | F16 | 411 MB |
v5-nano-retrieval-IQ1_M.gguf | IQ1_M | 97 MB |
v5-nano-retrieval-IQ1_S.gguf | IQ1_S | 95 MB |
v5-nano-retrieval-IQ2_M.gguf | IQ2_M | 108 MB |
v5-nano-retrieval-IQ2_XXS.gguf | IQ2_XXS | 101 MB |
v5-nano-retrieval-IQ4_NL.gguf | IQ4_NL | 145 MB |
v5-nano-retrieval-IQ4_XS.gguf | IQ4_XS | 142 MB |
v5-nano-retrieval-Q2_K.gguf | Q2_K | 124 MB |
v5-nano-retrieval-Q3_K_M.gguf | Q3_K_M | 137 MB |
v5-nano-retrieval-Q4_K_M.gguf | Q4_K_M | 150 MB |
v5-nano-retrieval-Q5_K_M.gguf | Q5_K_M | 161 MB |
v5-nano-retrieval-Q5_K_S.gguf | Q5_K_S | 159 MB |
v5-nano-retrieval-Q6_K.gguf | Q6_K | 173 MB |
v5-nano-retrieval-Q8_0.gguf | Q8_0 | 222 MB |
From the source README
jina-embeddings-v5-text-nano-retrieval-GGUF
GGUF quantizations of jina-embeddings-v5-text-nano-retrieval using llama.cpp. A 239M parameter multilingual embedding model quantized for efficient inference.
Elastic Inference Service | ArXiv | Blog
> [!IMPORTANT]
> We highly recommend to first read this blog post for more technical details and customized llama.cpp build.
Overview
`jina-embeddings-v5-text-nano-retrieval` is a task-specific embedding model for retrieval, part of the jina-embeddings-v5-text model family.
| Feature | Value |
| --- | --- |
| Parameters | 239M |
| Task | `retrieval` |
| Embedding Dimension | 768 |
| Matryoshka Dimensions | 32, 64, 128, 256, 512, 768 |
| Pooling Strategy | Last-token pooling |
| Base Model | jina-embeddings-v5-text-nano |
Usage with llama.cpp
via Elastic Inference Service
The fastest way to use v5-text in production. Elastic Inference Service (EIS) provides managed embedding inference with built-in scaling, so you can generate embeddings directly within your Elastic deployment.
PUT _inference/text_embedding/jina-v5
{
"service": "elastic",
"service_settings": {
"model_id": "jina-embeddings-v5-text-nano"
}
}
See the Elastic Inference Service documentation for setup details.
# Build llama.cpp (upstream)
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp…Card id model:hf:jinaai/jina-embeddings-v5-text-nano-retrieval-GGUF · collected 2026-10-02 21:00 UTC · JSON