Store › model › Embeddings
LFM2.5-Embedding-350M-GGUF
by LiquidAI · source Hugging Face · updated 2026-06-22
other219 MB~1 GB RAMsource aliveunlabeled
LFM2.5-Embedding-350M is a dense bi-encoder for fast multilingual retrieval. It produces a single vector per document — the smallest, fastest index — for reliable cross-lingual search across 11 langu…
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:LiquidAI/LFM2.5-Embedding-350M-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/LiquidAI/LFM2.5-Embedding-350M-GGUF
- License: other (custom license: read it at the source before installing)
- Requirements: about 1 GB of RAM, 219 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp.
- Tags:
sentence-transformersggufliquidlfm2lfm2.5edgesentence-similarityfeature-extractionllama.cppenesdefritptar
Numbers
- 7,403 downloads on Hugging Face
- 39 likes
- license other
- 0.2 GB for LFM2.5-Embedding-350M-Q4_K_M.gguf
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 21:00 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
LFM2.5-Embedding-350M-BF16.gguf | BF16 | 679 MB |
LFM2.5-Embedding-350M-F16.gguf | F16 | 679 MB |
LFM2.5-Embedding-350M-Q4_0.gguf | Q4_0 | 209 MB |
LFM2.5-Embedding-350M-Q4_K_M.gguf | Q4_K_M | 219 MB |
LFM2.5-Embedding-350M-Q5_K_M.gguf | Q5_K_M | 248 MB |
LFM2.5-Embedding-350M-Q6_K.gguf | Q6_K | 280 MB |
LFM2.5-Embedding-350M-Q8_0.gguf | Q8_0 | 362 MB |
From the source README
Try LFM • Documentation • LEAP
LFM2.5-Embedding-350M
LFM2.5-Embedding-350M is a dense bi-encoder for fast multilingual retrieval. It produces a single vector per document — the smallest, fastest index — for reliable cross-lingual search across 11 languages.
- Best-in-class multilingual accuracy for a dense embedder of its size.
- Inference speed is on par with much smaller models, thanks to the efficient LFM2 backbone.
- You can use it as a drop-in replacement in your current RAG pipelines.
Find more information about LFM2.5-Embedding-350M in our blog post.
🏃 How to run
Example usage with llama.cpp:
Start llama-server
```bash
llama-server -hf LiquidAI/LFM2.5-Embedding-350M-GGUF --embeddings
```
Make requests to embed queries and documents, and rank by cosine similarity (note the asymmetric `query: ` / `document: ` prompt prefixes)
❯ uv run dense-retrieve.py
Score: -0.1783 | Q: What is panda? | D: hi
Score: 0.0511 | Q: What is panda? | D: it is a bear
Score: 0.5657 | Q: What is panda? | D: The giant panda (Ailuropoda melanoleuca), sometimes called a panda bear or simply panda, is a bear species endemic to China.
```
# /// script
# requires-python = ">=3.10"
# dependencies = ["numpy", "requests"]
# ///Card id model:hf:LiquidAI/LFM2.5-Embedding-350M-GGUF · collected 2026-10-02 21:00 UTC · JSON