LogiShell store Open app

Store › model › Embeddings

jina-embeddings-v5-text-nano-retrieval-GGUF

by jinaai · source Hugging Face · updated 2026-02-27

cc-by-nc-4.0150 MB~1 GB RAMsource aliveunlabeled

jina-embeddings-v5-text-nano-retrieval-GGUF

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:jinaai/jina-embeddings-v5-text-nano-retrieval-GGUF and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-02 21:00 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
v5-nano-retrieval-F16.ggufF16411 MB
v5-nano-retrieval-IQ1_M.ggufIQ1_M97 MB
v5-nano-retrieval-IQ1_S.ggufIQ1_S95 MB
v5-nano-retrieval-IQ2_M.ggufIQ2_M108 MB
v5-nano-retrieval-IQ2_XXS.ggufIQ2_XXS101 MB
v5-nano-retrieval-IQ4_NL.ggufIQ4_NL145 MB
v5-nano-retrieval-IQ4_XS.ggufIQ4_XS142 MB
v5-nano-retrieval-Q2_K.ggufQ2_K124 MB
v5-nano-retrieval-Q3_K_M.ggufQ3_K_M137 MB
v5-nano-retrieval-Q4_K_M.ggufQ4_K_M150 MB
v5-nano-retrieval-Q5_K_M.ggufQ5_K_M161 MB
v5-nano-retrieval-Q5_K_S.ggufQ5_K_S159 MB
v5-nano-retrieval-Q6_K.ggufQ6_K173 MB
v5-nano-retrieval-Q8_0.ggufQ8_0222 MB

From the source README

jina-embeddings-v5-text-nano-retrieval-GGUF

GGUF quantizations of jina-embeddings-v5-text-nano-retrieval using llama.cpp. A 239M parameter multilingual embedding model quantized for efficient inference.

Elastic Inference Service | ArXiv | Blog

> [!IMPORTANT]
> We highly recommend to first read this blog post for more technical details and customized llama.cpp build.

Overview

`jina-embeddings-v5-text-nano-retrieval` is a task-specific embedding model for retrieval, part of the jina-embeddings-v5-text model family.
| Feature | Value |
| --- | --- |
| Parameters | 239M |
| Task | `retrieval` |
| Embedding Dimension | 768 |
| Matryoshka Dimensions | 32, 64, 128, 256, 512, 768 |
| Pooling Strategy | Last-token pooling |
| Base Model | jina-embeddings-v5-text-nano |

Usage with llama.cpp

via Elastic Inference Service

The fastest way to use v5-text in production. Elastic Inference Service (EIS) provides managed embedding inference with built-in scaling, so you can generate embeddings directly within your Elastic deployment.

PUT _inference/text_embedding/jina-v5
{
  "service": "elastic",
  "service_settings": {
    "model_id": "jina-embeddings-v5-text-nano"
  }
}

See the Elastic Inference Service documentation for setup details.

# Build llama.cpp (upstream)
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp…

Card id model:hf:jinaai/jina-embeddings-v5-text-nano-retrieval-GGUF · collected 2026-10-02 21:00 UTC · JSON