{"v":1,"id":"model:hf:ChristianAzinn/gte-small-gguf","slug":"model-christianazinn-gte-small-gguf","kind":"model","category":"embedding","title":"gte-small-gguf","summary":"General Text Embeddings (GTE) model. Towards General Text Embeddings with Multi-stage Contrastive Learning","source":{"provider":"hf","ref":"ChristianAzinn/gte-small-gguf","url":"https://huggingface.co/ChristianAzinn/gte-small-gguf","rev":"240acca7b64619cd22093a380dc266c4122d99b2","fetchedAt":"2026-10-02T20:59:31.687Z","etag":"W/\"12bc-3uodKEzteuMoJ57DN7kmzpUEtd4\""},"author":{"name":"ChristianAzinn","url":"https://huggingface.co/ChristianAzinn"},"license":{"spdx":"mit","raw":"mit","open":true},"metrics":{"downloads":5929,"downloadsWeek":327838,"likes":5,"takenAt":"2026-10-02T20:59:31.687Z"},"tags":["sentence-transformers","gguf","sentence-similarity","Sentence Transformers","mteb","bert","feature-extraction","en","deploy:azure"],"pipeline":"feature-extraction","links":{"github":"ggml-org/llama.cpp","npm":"node-llama-cpp"},"updatedAt":"2024-04-07T22:27:01.000Z","collectedAt":"2026-10-02T20:59:31.687Z","review":{"numbers":["5,929 downloads on Hugging Face","5 likes","license mit","0.0 GB for gte-small.Q4_K_M.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:59:31.687Z","http":200},"description":"# gte-small-gguf\n\nModel creator: [thenlper](https://huggingface.co/thenlper)\n\nOriginal model: [gte-small](https://huggingface.co/thenlper/gte-small)\n\n## Original Description\n\nGeneral Text Embeddings (GTE) model. [Towards General Text Embeddings with Multi-stage Contrastive Learning](https://arxiv.org/abs/2308.03281)\n\nThe GTE models are trained by Alibaba DAMO Academy. They are mainly based on the BERT framework and currently offer three different sizes of models, including [GTE-large](https://huggingface.co/thenlper/gte-large), [GTE-base](https://huggingface.co/thenlper/gte-base), and [GTE-small](https://huggingface.co/thenlper/gte-small). The GTE models are trained on a large-scale corpus of relevance text pairs, covering a wide range of domains and scenarios. This enables the GTE models to be applied to various downstream tasks of text embeddings, including **information retrieval**, **semantic textual similarity**, **text reranking**, etc.\n\n## Description\n\nThis repo contains GGUF format files for the gte-small embedding model.\n\nThese files were converted and quantized with llama.cpp [PR 5500](https://github.com/ggerganov/llama.cpp/pull/5500), commit [34aa045de](https://github.com/ggerganov/llama.cpp/pull/5500/commits/34aa045de44271ff7ad42858c75739303b8dc6eb), on a consumer RTX 4090.\n\nThis model supports up to 512 tokens of context.\n\n## Compatibility\n\nThese files are compatible with [llama.cpp](https://github.com/ggerganov/llama.cpp) as of commit [4524290e8](https://github.com/ggerganov/llama.cpp/commit/4524290e87b8e107cc2b56e1251751546f4b9051), as well as [LM Studio](https://lmstudio.ai/) as of version 0.2.19.\n\n# Meta-information\n## Explanation of quantisation methods\n\n  Click to see details\nThe methods available are:\n* GGML_TYPE_Q2_K - \"type-1\" 2-bit quantization in super-blocks containing 16 blocks, each block having 16 weight. Block scales and mins are quantized with 4 bits. This ends up effectivel…\n\nSource: https://huggingface.co/ChristianAzinn/gte-small-gguf","install":{"kind":"model","hfId":"ChristianAzinn/gte-small-gguf","gated":false,"format":"gguf","files":[{"name":"gte-small.Q2_K.gguf","size":25250080,"quant":"Q2_K","sha256":"71bc9beaecd0a3c5f075b8959f84c4cdf6c27dbc39930b0ab4d7c443b9373bc6"},{"name":"gte-small.Q3_K_L.gguf","size":27738400,"quant":"Q3_K_L","sha256":"3b7da433caf1dc18a00d49511d29f50e17def98fb1d5d4d49c469ce8802b76f2"},{"name":"gte-small.Q3_K_M.gguf","size":26724640,"quant":"Q3_K_M","sha256":"70c313740b579db2afffb61b625d70a70f2619331539661c5e95a2bfa20b2bf3"},{"name":"gte-small.Q3_K_S.gguf","size":25250080,"quant":"Q3_K_S","sha256":"c79755bc45783cf2e8a25fbae721a644cb2c177b1bee0cc96afe9048de652b41"},{"name":"gte-small.Q4_0.gguf","size":26190112,"quant":"Q4_0","sha256":"cd96dfa36de559194177d9dda0d51c393e5933d56eaa9b8d2013adf09677f1c0"},{"name":"gte-small.Q4_K_M.gguf","size":29203744,"quant":"Q4_K_M","sha256":"2b330c1579bac032397b48f5aa92b7b5ab2b94d72cc43cd15925db3ffd03fd61"},{"name":"gte-small.Q4_K_S.gguf","size":28217632,"quant":"Q4_K_S","sha256":"427635ac612c74e9609743bdcd1bd41ac8b1009ecc82515fc86b9ca926a23dfa"},{"name":"gte-small.Q5_0.gguf","size":28844320,"quant":"Q5_0","sha256":"20bd6030de6233c2d005f5b26e1d7bf66eb898818a0c96acb90300d9a4621f30"},{"name":"gte-small.Q5_K_M.gguf","size":30475552,"quant":"Q5_K_M","sha256":"baff7241c5bc5246a92bd4311019430b3cd5299297c3d3595767bf5b68786d8d"},{"name":"gte-small.Q5_K_S.gguf","size":29729056,"quant":"Q5_K_S","sha256":"2a06f8ad9dcc747557174d535ccdd3b3e9aa15f78020615927d74ee5cbc1d96c"},{"name":"gte-small.Q6_K.gguf","size":35092768,"quant":"Q6_K","sha256":"ad68d196e49e16c9c06119f14c87ae36fcc97a461246ce13ff91a7cabb48b374"},{"name":"gte-small.Q8_0.gguf","size":36806944,"quant":"Q8_0","sha256":"10978226002dd2db19dc4828083f6af5cff7d442ae76cb8dd9c4b7a8ba6550a7"},{"name":"gte-small_fp16.gguf","size":67308128,"sha256":"6c3d85a9af8ef795854d28cd25fa14bbf1638243d6c094d6f4c673b50b69271d"},{"name":"gte-small_fp32.gguf","size":133609568,"sha256":"edf8351430d6d9aea41c768644b9e01b7c8f9f01f7176ded435be1dbe8190718"}],"totalBytes":550441024,"suggestedFile":"gte-small.Q4_K_M.gguf","requirements":{"ramGb":1,"diskBytes":29203744,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp"],"command":"lsh models install hf:ChristianAzinn/gte-small-gguf"}}