{"v":1,"id":"model:hf:gpustack/bge-m3-GGUF","slug":"model-gpustack-bge-m3-gguf","kind":"model","category":"embedding","title":"bge-m3-GGUF","summary":"Model creator: BAAI Original model: bge-m3 GGUF quantization: based on llama.cpp release 61408e7f","source":{"provider":"hf","ref":"gpustack/bge-m3-GGUF","url":"https://huggingface.co/gpustack/bge-m3-GGUF","rev":"2d48f1737679ad900d5c26c5aad5410e9c70fdca","fetchedAt":"2026-10-02T21:00:41.566Z","etag":"W/\"f87-C6UmWgBJ6xnvm+esayHjJ/1vDso\""},"author":{"name":"gpustack","url":"https://huggingface.co/gpustack"},"license":{"spdx":"mit","raw":"mit","open":true},"metrics":{"downloads":67257,"downloadsWeek":4402710,"likes":57,"takenAt":"2026-10-02T21:00:41.566Z"},"tags":["sentence-transformers","gguf","feature-extraction","sentence-similarity","text-embeddings-inference","endpoints_compatible","deploy:azure"],"pipeline":"sentence-similarity","links":{"github":"FlagOpen/FlagEmbedding","npm":"@huggingface/transformers"},"updatedAt":"2024-10-31T08:31:47.000Z","collectedAt":"2026-10-02T21:00:41.566Z","review":{"numbers":["67,257 downloads on Hugging Face","57 likes","license mit","0.4 GB for bge-m3-Q4_K_M.gguf","4,402,710 npm downloads a week for @huggingface/transformers","latest @huggingface/transformers@4.3.0"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T21:00:41.566Z","http":200},"description":"# bge-m3-GGUF\n\n**Model creator**: [BAAI](https://huggingface.co/BAAI)\n**Original model**: [bge-m3](https://huggingface.co/BAAI/bge-m3)\n**GGUF quantization**: based on llama.cpp release [61408e7f](https://github.com/ggerganov/llama.cpp/commit/61408e7fad082dc44a11c8a9f1398da4837aad44)\n\n---\n\nFor more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding\n\n# BGE-M3 ([paper](https://arxiv.org/pdf/2402.03216.pdf), [code](https://github.com/FlagOpen/FlagEmbedding/tree/master/FlagEmbedding/BGE_M3))\n\nIn this project, we introduce BGE-M3, which is distinguished for its versatility in Multi-Functionality, Multi-Linguality, and Multi-Granularity.\n- Multi-Functionality: It can simultaneously perform the three common retrieval functionalities of embedding model: dense retrieval, multi-vector retrieval, and sparse retrieval.\n- Multi-Linguality: It can support more than 100 working languages.\n- Multi-Granularity: It is able to process inputs of different granularities, spanning from short sentences to long documents of up to 8192 tokens.\n\n**Some suggestions for retrieval pipeline in RAG**\n\nWe recommend to use the following pipeline: hybrid retrieval + re-ranking.\n- Hybrid retrieval leverages the strengths of various methods, offering higher accuracy and stronger generalization capabilities.\nA classic example: using both embedding retrieval and the BM25 algorithm.\nNow, you can try to use BGE-M3, which supports both embedding and sparse retrieval.\nThis allows you to obtain token weights (similar to the BM25) without any additional cost when generate dense embeddings.\nTo use hybrid retrieval, you can refer to [Vespa](https://github.com/vespa-engine/pyvespa/blob/master/docs/sphinx/source/examples/mother-of-all-embedding-models-cloud.ipynb\n) and [Milvus](https://github.com/milvus-io/pymilvus/blob/master/examples/hello_hybrid_sparse_dense.py).\n\n- As cross-encoder models, re-ranker demonstrates higher accura…\n\nSource: https://huggingface.co/gpustack/bge-m3-GGUF","install":{"kind":"model","hfId":"gpustack/bge-m3-GGUF","gated":false,"format":"gguf","files":[{"name":"bge-m3-FP16.gguf","size":1157671200,"sha256":"daec91ffb5dd0c27411bd71f29932917c49cf529a641d0168496c3a501e3062c"},{"name":"bge-m3-Q2_K.gguf","size":366114880,"quant":"Q2_K","sha256":"8b0dacc25f20c7375700e7143e30ce48a8fbe44e66b035de243113b9cc5d44d8"},{"name":"bge-m3-Q3_K.gguf","size":402290752,"quant":"Q3_K","sha256":"c4e95d4c9ad1680bdd44735e36e10b68f71a9470572fe3b1b5d9fc147299b3e1"},{"name":"bge-m3-Q4_0.gguf","size":421558336,"quant":"Q4_0","sha256":"6eaafd7b20eecbcf9d239c670b8a06a0c23ce3ef5b3bbf37ec02db327203ac02"},{"name":"bge-m3-Q4_K_M.gguf","size":437778496,"quant":"Q4_K_M","sha256":"6d39681b26c61279ac1f82db35a04a05009e94c415b51c858ff571489a82fc06"},{"name":"bge-m3-Q5_0.gguf","size":459307072,"quant":"Q5_0","sha256":"a2b6d898a86d9c702b97087806e6a3371b154d915bcbbb0d2ba47b00ca734fe2"},{"name":"bge-m3-Q5_K_M.gguf","size":467662912,"quant":"Q5_K_M","sha256":"f93897db57c4385f1cde3f59234bececa234c0f9bafc646dfdaeebe7f65ea84d"},{"name":"bge-m3-Q6_K.gguf","size":499415104,"quant":"Q6_K","sha256":"d7b2ddf8880de940dbdf484c2d1c0b2db5b68f0c342634819cdb3d32e5703c0a"},{"name":"bge-m3-Q8_0.gguf","size":634553760,"quant":"Q8_0","sha256":"950f4a8e5e19477a6d3c26d2f162233c20002c601f75e4b002e3239997821167"}],"totalBytes":4846352512,"suggestedFile":"bge-m3-Q4_K_M.gguf","requirements":{"ramGb":1,"diskBytes":437778496,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp"],"command":"lsh models install hf:gpustack/bge-m3-GGUF"}}