Store › model › Embeddings
embeddinggemma-300M-qat-q4_0-GGUF
by ggml-org · source Hugging Face · updated 2025-09-15
gemma265 MB~1 GB RAMsource aliveunlabeled
Alternatively, the llama-embedding command line tool can be used: sh llama-embedding -hf ggml-org/embeddinggemma-300M-qat-q40-GGUF --verbose-prompt -p "Hello embeddings"
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:ggml-org/embeddinggemma-300M-qat-q4_0-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/ggml-org/embeddinggemma-300M-qat-q4_0-GGUF
- License: gemma (restricted license: terms at the source apply)
- Requirements: about 1 GB of RAM, 265 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp.
- Tags:
sentence-transformersggufsentence-similarityfeature-extractionendpoints_compatible
Numbers
- 13,751 downloads on Hugging Face
- 6 likes
- license gemma
- 0.3 GB for embeddinggemma-300M-qat-Q4_0.gguf
- 327,838 npm downloads a week for node-llama-cpp
- latest node-llama-cpp@3.22.1
Numbers as of 2026-10-02 20:59 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
embeddinggemma-300M-qat-Q4_0.gguf | Q4_0 | 265 MB |
From the source README
embeddinggemma-300M-qat-q4_0 GGUF
Recommended way to run this model:
llama-server -hf ggml-org/embeddinggemma-300M-qat-q4_0-GGUF --embeddings
Then the endpoint can be accessed at http://localhost:8080/embedding, for
example using `curl`:
```console
curl --request POST \
--url http://localhost:8080/embedding \
--header "Content-Type: application/json" \
--data '{"input": "Hello embeddings"}' \
--silent
```
Alternatively, the `llama-embedding` command line tool can be used:
```sh
llama-embedding -hf ggml-org/embeddinggemma-300M-qat-q4_0-GGUF --verbose-prompt -p "Hello embeddings"
```
#### embd_normalize
When a model uses pooling, or the pooling method is specified using `--pooling`,
the normalization can be controlled by the `embd_normalize` parameter.
The default value is `2` which means that the embeddings are normalized using
the Euclidean norm (L2). Other options are:
* -1 No normalization
* 0 Max absolute
* 1 Taxicab
* 2 Euclidean/L2
* \>2 P-Norm
This can be passed in the request body to `llama-server`, for example:
```sh
--data '{"input": "Hello embeddings", "embd_normalize": -1}' \
```
And for `llama-embedding`, by passing `--embd-normalize `, for example:
```sh
llama-embedding -hf ggml-org/embeddinggemma-300M-qat-q4_0-GGUF --embd-normalize -1 -p "Hello embeddings"
```
Source: https://huggingface.co/ggml-org/embeddinggemma-300M-qat-q4_0-GGUF
Card id model:hf:ggml-org/embeddinggemma-300M-qat-q4_0-GGUF · collected 2026-10-02 20:59 UTC · JSON