Store › model › Embeddings
bge-m3-GGUF
by gpustack · source Hugging Face · updated 2024-10-31
mit417 MB~1 GB RAMsource aliveunlabeled
Model creator: BAAI Original model: bge-m3 GGUF quantization: based on llama.cpp release 61408e7f
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:gpustack/bge-m3-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/gpustack/bge-m3-GGUF
- License: mit
- Requirements: about 1 GB of RAM, 417 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp.
- Tags:
sentence-transformersgguffeature-extractionsentence-similaritytext-embeddings-inferenceendpoints_compatibledeploy:azure
Numbers
- 67,257 downloads on Hugging Face
- 57 likes
- license mit
- 0.4 GB for bge-m3-Q4_K_M.gguf
- 4,402,710 npm downloads a week for @huggingface/transformers
- latest @huggingface/transformers@4.3.0
Numbers as of 2026-10-02 21:00 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
bge-m3-FP16.gguf | 1.1 GB | |
bge-m3-Q2_K.gguf | Q2_K | 349 MB |
bge-m3-Q3_K.gguf | Q3_K | 384 MB |
bge-m3-Q4_0.gguf | Q4_0 | 402 MB |
bge-m3-Q4_K_M.gguf | Q4_K_M | 417 MB |
bge-m3-Q5_0.gguf | Q5_0 | 438 MB |
bge-m3-Q5_K_M.gguf | Q5_K_M | 446 MB |
bge-m3-Q6_K.gguf | Q6_K | 476 MB |
bge-m3-Q8_0.gguf | Q8_0 | 605 MB |
From the source README
bge-m3-GGUF
Model creator: BAAI
Original model: bge-m3
GGUF quantization: based on llama.cpp release 61408e7f
For more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding
BGE-M3 (paper, code)
In this project, we introduce BGE-M3, which is distinguished for its versatility in Multi-Functionality, Multi-Linguality, and Multi-Granularity.
- Multi-Functionality: It can simultaneously perform the three common retrieval functionalities of embedding model: dense retrieval, multi-vector retrieval, and sparse retrieval.
- Multi-Linguality: It can support more than 100 working languages.
- Multi-Granularity: It is able to process inputs of different granularities, spanning from short sentences to long documents of up to 8192 tokens.
Some suggestions for retrieval pipeline in RAG
We recommend to use the following pipeline: hybrid retrieval + re-ranking.
- Hybrid retrieval leverages the strengths of various methods, offering higher accuracy and stronger generalization capabilities.
A classic example: using both embedding retrieval and the BM25 algorithm.
Now, you can try to use BGE-M3, which supports both embedding and sparse retrieval.
This allows you to obtain token weights (similar to the BM25) without any additional cost when generate dense embeddings.
To use hybrid retrieval, you can refer to Vespa and Milvus.
- As cross-encoder models, re-ranker demonstrates higher accura…
Source: https://huggingface.co/gpustack/bge-m3-GGUF
Card id model:hf:gpustack/bge-m3-GGUF · collected 2026-10-02 21:00 UTC · JSON