{"v":1,"id":"model:hf:ggml-org/embeddinggemma-300m-qat-q8_0-GGUF","slug":"model-ggml-org-embeddinggemma-300m-qat-q8-0-gguf","kind":"model","category":"embedding","title":"embeddinggemma-300m-qat-q8_0-GGUF","summary":"Alternatively, the llama-embedding command line tool can be used: sh llama-embedding -hf ggml-org/embeddinggemma-300m-qat-q80-GGUF --verbose-prompt -p \"Hello embeddings\"","source":{"provider":"hf","ref":"ggml-org/embeddinggemma-300m-qat-q8_0-GGUF","url":"https://huggingface.co/ggml-org/embeddinggemma-300m-qat-q8_0-GGUF","rev":"66f974f8cd48cc3b9c41c516b95508e75b4bee64","fetchedAt":"2026-10-02T21:01:13.620Z","etag":"W/\"622-eF8OIiJ0fPHVYvDfbDDgwRvAmSI\""},"author":{"name":"ggml-org","url":"https://huggingface.co/ggml-org"},"license":{"spdx":"gemma","raw":"gemma","open":false,"note":"restricted license: terms at the source apply"},"metrics":{"downloads":39751,"downloadsWeek":327838,"likes":21,"takenAt":"2026-10-02T21:01:13.620Z"},"tags":["sentence-transformers","gguf","sentence-similarity","feature-extraction","endpoints_compatible"],"pipeline":"feature-extraction","links":{"github":"google-deepmind/gemma","npm":"node-llama-cpp"},"updatedAt":"2025-09-15T07:52:04.000Z","collectedAt":"2026-10-02T21:01:13.620Z","review":{"numbers":["39,751 downloads on Hugging Face","21 likes","license gemma","0.3 GB for embeddinggemma-300m-qat-Q8_0.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T21:01:13.620Z","http":200},"description":"# embeddinggemma-300m-qat-q8_0 GGUF\n\nRecommended way to run this model:\n\n```sh\nllama-server -hf ggml-org/embeddinggemma-300m-qat-q8_0-GGUF --embeddings\n```\n\nThen the endpoint can be accessed at http://localhost:8080/embedding, for\nexample using `curl`:\n```console\ncurl --request POST \\\n    --url http://localhost:8080/embedding \\\n    --header \"Content-Type: application/json\" \\\n    --data '{\"input\": \"Hello embeddings\"}' \\\n    --silent\n```\n\nAlternatively, the `llama-embedding` command line tool can be used:\n```sh\nllama-embedding -hf ggml-org/embeddinggemma-300m-qat-q8_0-GGUF --verbose-prompt -p \"Hello embeddings\"\n```\n\n#### embd_normalize\nWhen a model uses pooling, or the pooling method is specified using `--pooling`,\nthe normalization can be controlled by the `embd_normalize` parameter.\n\nThe default value is `2` which means that the embeddings are normalized using\nthe Euclidean norm (L2). Other options are:\n* -1 No normalization\n*  0 Max absolute\n*  1 Taxicab\n*  2 Euclidean/L2\n* \\>2 P-Norm\n\nThis can be passed in the request body to `llama-server`, for example:\n```sh\n    --data '{\"input\": \"Hello embeddings\", \"embd_normalize\": -1}' \\\n```\n\nAnd for `llama-embedding`, by passing `--embd-normalize `, for example:\n```sh\nllama-embedding -hf ggml-org/embeddinggemma-300m-qat-q8_0-GGUF  --embd-normalize -1 -p \"Hello embeddings\"\n```\n\nSource: https://huggingface.co/ggml-org/embeddinggemma-300m-qat-q8_0-GGUF","install":{"kind":"model","hfId":"ggml-org/embeddinggemma-300m-qat-q8_0-GGUF","gated":false,"format":"gguf","files":[{"name":"embeddinggemma-300m-qat-Q8_0.gguf","size":328577056,"quant":"Q8_0","sha256":"6fa0c02a9c302be6f977521d399b4de3a46310a4f2621ee0063747881b673f67"}],"totalBytes":328577056,"suggestedFile":"embeddinggemma-300m-qat-Q8_0.gguf","requirements":{"ramGb":1,"diskBytes":328577056,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp"],"command":"lsh models install hf:ggml-org/embeddinggemma-300m-qat-q8_0-GGUF"}}