{"v":1,"id":"model:hf:LiquidAI/LFM2.5-Embedding-350M-GGUF","slug":"model-liquidai-lfm2-5-embedding-350m-gguf","kind":"model","category":"embedding","title":"LFM2.5-Embedding-350M-GGUF","summary":"LFM2.5-Embedding-350M is a dense bi-encoder for fast multilingual retrieval. It produces a single vector per document — the smallest, fastest index — for reliable cross-lingual search across 11 langu…","source":{"provider":"hf","ref":"LiquidAI/LFM2.5-Embedding-350M-GGUF","url":"https://huggingface.co/LiquidAI/LFM2.5-Embedding-350M-GGUF","rev":"a80de9c5b941d429104f0038292a0ef5a860e486","fetchedAt":"2026-10-02T21:00:54.720Z","etag":"W/\"133a-z9F6ZBkB1ZtQK17L4ePFRurMhzo\""},"author":{"name":"LiquidAI","url":"https://huggingface.co/LiquidAI"},"license":{"spdx":null,"raw":"other","open":null,"note":"custom license: read it at the source before installing"},"metrics":{"downloads":7403,"downloadsWeek":327838,"likes":39,"takenAt":"2026-10-02T21:00:54.720Z"},"tags":["sentence-transformers","gguf","liquid","lfm2","lfm2.5","edge","sentence-similarity","feature-extraction","llama.cpp","en","es","de","fr","it","pt","ar","sv","no","ja","ko","endpoints_compatible","conversational"],"pipeline":"sentence-similarity","links":{"github":"ggml-org/llama.cpp","npm":"node-llama-cpp"},"updatedAt":"2026-06-22T17:55:00.000Z","collectedAt":"2026-10-02T21:00:54.720Z","review":{"numbers":["7,403 downloads on Hugging Face","39 likes","license other","0.2 GB for LFM2.5-Embedding-350M-Q4_K_M.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T21:00:54.720Z","http":200},"description":"Try LFM • Documentation • LEAP\n\n# LFM2.5-Embedding-350M\n\nLFM2.5-Embedding-350M is a dense bi-encoder for fast multilingual retrieval. It produces a single vector per document — the smallest, fastest index — for reliable cross-lingual search across 11 languages.\n\n- **Best-in-class multilingual accuracy** for a dense embedder of its size.\n- Inference speed is **on par with much smaller models**, thanks to the efficient LFM2 backbone.\n- You can use it as a **drop-in replacement** in your current RAG pipelines.\n\nFind more information about LFM2.5-Embedding-350M in our [blog post](https://liquid-ai-v3-c7c6d49467ac-bf50aea57dc57.webflow.io/blog/lfm2-5-retrievers).\n\n## 🏃 How to run\n\nExample usage with [llama.cpp](https://github.com/ggml-org/llama.cpp):\n\nStart llama-server\n```bash\nllama-server -hf LiquidAI/LFM2.5-Embedding-350M-GGUF --embeddings\n```\n\nMake requests to embed queries and documents, and rank by cosine similarity (note the asymmetric `query: ` / `document: ` prompt prefixes)\n\n```bash\n❯ uv run dense-retrieve.py\n\nScore: -0.1783 | Q: What is panda? | D: hi\nScore:  0.0511 | Q: What is panda? | D: it is a bear\nScore:  0.5657 | Q: What is panda? | D: The giant panda (Ailuropoda melanoleuca), sometimes called a panda bear or simply panda, is a bear species endemic to China.\n```\n\n```python\n# /// script\n# requires-python = \">=3.10\"\n# dependencies = [\"numpy\", \"requests\"]\n# ///\n\n# dense-retrieve.py\nimport numpy as np, requests\n\nQUERY_PREFIX, DOC_PREFIX = \"query: \", \"document: \"\n\ndef embed(text: str) -> np.ndarray:\n    r = requests.post(\n        \"http://localhost:8080/v1/embeddings\",\n        json={\"input\": text},\n    )\n    v = np.array(r.json()[\"data\"][0][\"embedding\"])\n    return v / np.linalg.norm(v)\n\ndocs = [\n    \"hi\",\n    \"it is a bear\",\n    \"The giant panda (Ailuropoda melanoleuca), sometimes called a panda bear or simply panda, is a bear species endemic to China.\",\n]\nquery = \"What is panda?\"\n\nq = emb…\n\nSource: https://huggingface.co/LiquidAI/LFM2.5-Embedding-350M-GGUF","install":{"kind":"model","hfId":"LiquidAI/LFM2.5-Embedding-350M-GGUF","gated":false,"format":"gguf","files":[{"name":"LFM2.5-Embedding-350M-BF16.gguf","size":711484160,"quant":"BF16","sha256":"c01a8eae5fcc937f098a84330419d8fc9da6322b2bb598572273fd2294f2eb1e"},{"name":"LFM2.5-Embedding-350M-F16.gguf","size":711484160,"quant":"F16","sha256":"da715f5bbe2a91dd518ba92912ad3502bf6a06ecad1397a12719e4359743dc54"},{"name":"LFM2.5-Embedding-350M-Q4_0.gguf","size":219308800,"quant":"Q4_0","sha256":"08bb9bfdfb516146b7aeb6d51994206b38e8b31c9e6e66a7bdcebbbc0afa3cea"},{"name":"LFM2.5-Embedding-350M-Q4_K_M.gguf","size":229311232,"quant":"Q4_K_M","sha256":"4d7aa9dc6406a10fc3dec2c11f8f06781af063bf49211b8e4132e9b876d3f32a"},{"name":"LFM2.5-Embedding-350M-Q5_K_M.gguf","size":260375296,"quant":"Q5_K_M","sha256":"c65f59f04e1def36f83af061706a0df71c6203671e121fe087f99854438122c1"},{"name":"LFM2.5-Embedding-350M-Q6_K.gguf","size":293380864,"quant":"Q6_K","sha256":"7055c76bbbad27760862c7719a341ac8ae73363e2092450f312d3a86d2a1fef0"},{"name":"LFM2.5-Embedding-350M-Q8_0.gguf","size":379216640,"quant":"Q8_0","sha256":"6ec5f8e8750dbc8a0e40c431fd1b7b07a13688136b2244c5a1364b54d9032599"}],"totalBytes":2804561152,"suggestedFile":"LFM2.5-Embedding-350M-Q4_K_M.gguf","requirements":{"ramGb":1,"diskBytes":229311232,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp"],"command":"lsh models install hf:LiquidAI/LFM2.5-Embedding-350M-GGUF"}}