LogiShell store Open app

Store › model › Embeddings

bge-m3-GGUF

by gpustack · source Hugging Face · updated 2024-10-31

mit417 MB~1 GB RAMsource aliveunlabeled

Model creator: BAAI Original model: bge-m3 GGUF quantization: based on llama.cpp release 61408e7f

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:gpustack/bge-m3-GGUF and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-02 21:00 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
bge-m3-FP16.gguf1.1 GB
bge-m3-Q2_K.ggufQ2_K349 MB
bge-m3-Q3_K.ggufQ3_K384 MB
bge-m3-Q4_0.ggufQ4_0402 MB
bge-m3-Q4_K_M.ggufQ4_K_M417 MB
bge-m3-Q5_0.ggufQ5_0438 MB
bge-m3-Q5_K_M.ggufQ5_K_M446 MB
bge-m3-Q6_K.ggufQ6_K476 MB
bge-m3-Q8_0.ggufQ8_0605 MB

From the source README

bge-m3-GGUF

Model creator: BAAI
Original model: bge-m3
GGUF quantization: based on llama.cpp release 61408e7f

For more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding

BGE-M3 (paper, code)

In this project, we introduce BGE-M3, which is distinguished for its versatility in Multi-Functionality, Multi-Linguality, and Multi-Granularity.
- Multi-Functionality: It can simultaneously perform the three common retrieval functionalities of embedding model: dense retrieval, multi-vector retrieval, and sparse retrieval.
- Multi-Linguality: It can support more than 100 working languages.
- Multi-Granularity: It is able to process inputs of different granularities, spanning from short sentences to long documents of up to 8192 tokens.

Some suggestions for retrieval pipeline in RAG

We recommend to use the following pipeline: hybrid retrieval + re-ranking.
- Hybrid retrieval leverages the strengths of various methods, offering higher accuracy and stronger generalization capabilities.
A classic example: using both embedding retrieval and the BM25 algorithm.
Now, you can try to use BGE-M3, which supports both embedding and sparse retrieval.
This allows you to obtain token weights (similar to the BM25) without any additional cost when generate dense embeddings.
To use hybrid retrieval, you can refer to Vespa and Milvus.

  • As cross-encoder models, re-ranker demonstrates higher accura…

Source: https://huggingface.co/gpustack/bge-m3-GGUF

Card id model:hf:gpustack/bge-m3-GGUF · collected 2026-10-02 21:00 UTC · JSON