{"v":1,"id":"model:hf:ggml-org/embeddinggemma-300M-qat-q4_0-GGUF","slug":"model-ggml-org-embeddinggemma-300m-qat-q4-0-gguf","kind":"model","category":"embedding","title":"embeddinggemma-300M-qat-q4_0-GGUF","summary":"Alternatively, the llama-embedding command line tool can be used: sh llama-embedding -hf ggml-org/embeddinggemma-300M-qat-q40-GGUF --verbose-prompt -p \"Hello embeddings\"","source":{"provider":"hf","ref":"ggml-org/embeddinggemma-300M-qat-q4_0-GGUF","url":"https://huggingface.co/ggml-org/embeddinggemma-300M-qat-q4_0-GGUF","rev":"8dd0ca2a66a8f14470acb0e2a71f801afbc5fb73","fetchedAt":"2026-10-02T20:59:25.624Z","etag":"W/\"5f8-69m4bFc37lmFj2k8jDxBAcsxk3c\""},"author":{"name":"ggml-org","url":"https://huggingface.co/ggml-org"},"license":{"spdx":"gemma","raw":"gemma","open":false,"note":"restricted license: terms at the source apply"},"metrics":{"downloads":13751,"downloadsWeek":327838,"likes":6,"takenAt":"2026-10-02T20:59:25.624Z"},"tags":["sentence-transformers","gguf","sentence-similarity","feature-extraction","endpoints_compatible"],"pipeline":"feature-extraction","links":{"github":"google-deepmind/gemma","npm":"node-llama-cpp"},"updatedAt":"2025-09-15T07:51:38.000Z","collectedAt":"2026-10-02T20:59:25.624Z","review":{"numbers":["13,751 downloads on Hugging Face","6 likes","license gemma","0.3 GB for embeddinggemma-300M-qat-Q4_0.gguf","327,838 npm downloads a week for node-llama-cpp","latest node-llama-cpp@3.22.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:59:25.624Z","http":200},"description":"# embeddinggemma-300M-qat-q4_0 GGUF\n\nRecommended way to run this model:\n\n```sh\nllama-server -hf ggml-org/embeddinggemma-300M-qat-q4_0-GGUF --embeddings\n```\n\nThen the endpoint can be accessed at http://localhost:8080/embedding, for\nexample using `curl`:\n```console\ncurl --request POST \\\n    --url http://localhost:8080/embedding \\\n    --header \"Content-Type: application/json\" \\\n    --data '{\"input\": \"Hello embeddings\"}' \\\n    --silent\n```\n\nAlternatively, the `llama-embedding` command line tool can be used:\n```sh\nllama-embedding -hf ggml-org/embeddinggemma-300M-qat-q4_0-GGUF --verbose-prompt -p \"Hello embeddings\"\n```\n\n#### embd_normalize\nWhen a model uses pooling, or the pooling method is specified using `--pooling`,\nthe normalization can be controlled by the `embd_normalize` parameter.\n\nThe default value is `2` which means that the embeddings are normalized using\nthe Euclidean norm (L2). Other options are:\n* -1 No normalization\n*  0 Max absolute\n*  1 Taxicab\n*  2 Euclidean/L2\n* \\>2 P-Norm\n\nThis can be passed in the request body to `llama-server`, for example:\n```sh\n    --data '{\"input\": \"Hello embeddings\", \"embd_normalize\": -1}' \\\n```\n\nAnd for `llama-embedding`, by passing `--embd-normalize `, for example:\n```sh\nllama-embedding -hf ggml-org/embeddinggemma-300M-qat-q4_0-GGUF  --embd-normalize -1 -p \"Hello embeddings\"\n```\n\nSource: https://huggingface.co/ggml-org/embeddinggemma-300M-qat-q4_0-GGUF","install":{"kind":"model","hfId":"ggml-org/embeddinggemma-300M-qat-q4_0-GGUF","gated":false,"format":"gguf","files":[{"name":"embeddinggemma-300M-qat-Q4_0.gguf","size":277852192,"quant":"Q4_0","sha256":"50d28e22432a148f6f8a86eab3700f92add5d1f54baf7790675a2a4dadbccf26"}],"totalBytes":277852192,"suggestedFile":"embeddinggemma-300M-qat-Q4_0.gguf","requirements":{"ramGb":1,"diskBytes":277852192,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp"],"command":"lsh models install hf:ggml-org/embeddinggemma-300M-qat-q4_0-GGUF"}}