{"v":1,"id":"model:hf:ChristianAzinn/gte-large-gguf","slug":"model-christianazinn-gte-large-gguf","kind":"model","category":"embedding","title":"gte-large-gguf","summary":"General Text Embeddings (GTE) model. Towards General Text Embeddings with Multi-stage Contrastive Learning","source":{"provider":"hf","ref":"ChristianAzinn/gte-large-gguf","url":"https://huggingface.co/ChristianAzinn/gte-large-gguf","rev":"f9fa5479908e72c2a8b9d6ba112911cd1e51be53","fetchedAt":"2026-10-02T20:59:49.547Z","etag":"W/\"12db-Tkz5QTMyyPxN2hE4Oh3WaRyGEws\""},"author":{"name":"ChristianAzinn","url":"https://huggingface.co/ChristianAzinn"},"license":{"spdx":"mit","raw":"mit","open":true},"metrics":{"downloads":2891,"downloadsWeek":327838,"likes":1,"takenAt":"2026-10-02T20:59:49.547Z"},"tags":["sentence-transformers","gguf","mteb","bert","sentence-similarity","Sentence Transformers","feature-extraction","en","deploy:azure"],"pipeline":"feature-extraction","links":{"github":"ggml-org/llama.cpp","npm":"node-llama-cpp"},"updatedAt":"2024-04-07T22:16:08.000Z","collectedAt":"2026-10-02T20:59:49.547Z","review":{"numbers":["2,891 downloads on Hugging Face","1 likes","license mit","0.2 GB for gte-large.Q4_K_M.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:59:49.547Z","http":200},"description":"# gte-large-gguf\n\nModel creator: [thenlper](https://huggingface.co/thenlper)\n\nOriginal model: [gte-large](https://huggingface.co/thenlper/gte-large)\n\n## Original Description\n\nGeneral Text Embeddings (GTE) model. [Towards General Text Embeddings with Multi-stage Contrastive Learning](https://arxiv.org/abs/2308.03281)\n\nThe GTE models are trained by Alibaba DAMO Academy. They are mainly based on the BERT framework and currently offer three different sizes of models, including [GTE-large](https://huggingface.co/thenlper/gte-large), [GTE-base](https://huggingface.co/thenlper/gte-base), and [GTE-small](https://huggingface.co/thenlper/gte-small). The GTE models are trained on a large-scale corpus of relevance text pairs, covering a wide range of domains and scenarios. This enables the GTE models to be applied to various downstream tasks of text embeddings, including **information retrieval**, **semantic textual similarity**, **text reranking**, etc.\n\n## Description\n\nThis repo contains GGUF format files for the gte-large embedding model.\n\nThese files were converted and quantized with llama.cpp [PR 5500](https://github.com/ggerganov/llama.cpp/pull/5500), commit [34aa045de](https://github.com/ggerganov/llama.cpp/pull/5500/commits/34aa045de44271ff7ad42858c75739303b8dc6eb), on a consumer RTX 4090.\n\nThis model supports up to 512 tokens of context.\n\n## Compatibility\n\nThese files are compatible with [llama.cpp](https://github.com/ggerganov/llama.cpp) as of commit [4524290e8](https://github.com/ggerganov/llama.cpp/commit/4524290e87b8e107cc2b56e1251751546f4b9051), as well as [LM Studio](https://lmstudio.ai/) as of version 0.2.19.\n\n# Meta-information\n## Explanation of quantisation methods\n\n  Click to see details\nThe methods available are:\n* GGML_TYPE_Q2_K - \"type-1\" 2-bit quantization in super-blocks containing 16 blocks, each block having 16 weight. Block scales and mins are quantized with 4 bits. This ends up effectivel…\n\nSource: https://huggingface.co/ChristianAzinn/gte-large-gguf","install":{"kind":"model","hfId":"ChristianAzinn/gte-large-gguf","gated":false,"format":"gguf","files":[{"name":"gte-large.Q2_K.gguf","size":144227872,"quant":"Q2_K","sha256":"32b6d1ad23a5440367b4b6369e37c6f3947171160a1f82ed336508c597a743ea"},{"name":"gte-large.Q3_K_L.gguf","size":198491680,"quant":"Q3_K_L","sha256":"7e3f79f72776b52b794a68eba3ead5627c2548abf0cc848f80ea59486b2f6ca3"},{"name":"gte-large.Q3_K_M.gguf","size":181452320,"quant":"Q3_K_M","sha256":"3556efd12404b49b3bdcd8b7f5d23e05fe4a7ede50a3d62980ca77a5fe561a82"},{"name":"gte-large.Q3_K_S.gguf","size":159563296,"quant":"Q3_K_S","sha256":"5e7ade73825a6d4b57b58f9de2343997c201c660c7fdafb3a430d1da98197dc9"},{"name":"gte-large.Q4_0.gguf","size":199671328,"quant":"Q4_0","sha256":"f2b99b0da24169193da8ca1a13c3169c0e16e6b85288613277c93f55fb0c49d6"},{"name":"gte-large.Q4_K_M.gguf","size":215891488,"quant":"Q4_K_M","sha256":"8dd25bb26d40ee8e11200a6b94a15d03403f055d04f3c07f5281d071a3d65a4d"},{"name":"gte-large.Q4_K_S.gguf","size":203341344,"quant":"Q4_K_S","sha256":"9048024b1161152f5f7dbeff7764a3218b2a0fbeff0a2e0bb9efa980143c54c3"},{"name":"gte-large.Q5_0.gguf","size":237420064,"quant":"Q5_0","sha256":"e708a09159eb3d5419715361691517ae77e76adefc2e4e1fe09f8f98c97a78b1"},{"name":"gte-large.Q5_K_M.gguf","size":245775904,"quant":"Q5_K_M","sha256":"2ce3133f45be3a84a83e2c994ee24167e308416b8ceb05cdaa836f1ae924a986"},{"name":"gte-large.Q5_K_S.gguf","size":237420064,"quant":"Q5_K_S","sha256":"d79003e24735f1aa6bcd8552743f0a6d0121dd774a485751abcff3a76ccb32a0"},{"name":"gte-large.Q6_K.gguf","size":277528096,"quant":"Q6_K","sha256":"2d0cae1850945f716b49a9cf19c7f1f6136ff5f1f6f658c0720f7288d48d5e45"},{"name":"gte-large.Q8_0.gguf","size":358235712,"quant":"Q8_0","sha256":"d6412ebf17908239458ce1a4fb4dd8b260dd4f2dcf164f50f53f3b0b971ade3b"},{"name":"gte-large_fp16.gguf","size":669603712,"sha256":"939f1fb3fcc70f2a250a7e7ad7c2fbdc1397d46f9a8055d053e451829c5293fb"},{"name":"gte-large_fp32.gguf","size":1337141120,"sha256":"e1d8caa7970221a615933d5e6201d9df146dea15694bd259b548b97d7fc0d63b"}],"totalBytes":4665764000,"suggestedFile":"gte-large.Q4_K_M.gguf","requirements":{"ramGb":1,"diskBytes":215891488,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp"],"command":"lsh models install hf:ChristianAzinn/gte-large-gguf"}}