Store › model › Embeddings
Nemotron-3-Embed-1B-Q4_K_M-GGUF
by zenmagnets · source Hugging Face · updated 2026-07-17
openmdw-1.1715 MB~2 GB RAMsource aliveunlabeled
An independently converted and quantized GGUF of NVIDIA's nvidia/Nemotron-3-Embed-1B-BF16, prepared for local embedding inference in LM Studio and llama.cpp-compatible runtimes.
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:zenmagnets/Nemotron-3-Embed-1B-Q4_K_M-GGUF and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/zenmagnets/Nemotron-3-Embed-1B-Q4_K_M-GGUF
- License: openmdw-1.1
- Requirements: about 2 GB of RAM, 715 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with llama.cpp.
- Tags:
ggufQ4_K_Mtext-embeddingsfeature-extractionretrievalsemantic-searchraglm-studiollama-cppsentence-similaritymultilingualenarasbnbg
Numbers
- 102,925 downloads on Hugging Face
- 5 likes
- license openmdw-1.1
- 0.7 GB for nemotron-3-embed-1b-q4_k_m.gguf
- 18,539 stars on NVIDIA-NeMo/Speech
- 325 open issues and PRs
- last release v3.0.0 on 2026-08-07
Numbers as of 2026-10-02 20:53 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
nemotron-3-embed-1b-q4_k_m.gguf | Q4_K_M | 715 MB |
From the source README
Nemotron-3-Embed-1B Q4_K_M GGUF
An independently converted and quantized GGUF of NVIDIA's
`nvidia/Nemotron-3-Embed-1B-BF16`,
prepared for local embedding inference in LM Studio and llama.cpp-compatible runtimes.
This repository is not an official NVIDIA release and is not affiliated with or endorsed by NVIDIA.
File
| File | Quantization | Size | SHA-256 |
|---|---:|---:|---|
| `nemotron-3-embed-1b-q4_k_m.gguf` | Q4_K_M | 749,352,096 bytes (714.6 MiB) | `9a74166f51dbc280073748fa199bea49283bd21f7f9280f2dec2b4d975ddfd1d` |
The model produces 2,048-dimensional, L2-normalized embeddings. Its GGUF metadata declares a
262,144-token maximum context. The release was functionally tested at a 4,096-token context; very
large contexts were not validated and require substantially more memory.
Use with LM Studio
Download the GGUF from Files and versions, then drag it into LM Studio or place it in LM Studio's
models directory. LM Studio should classify it as an embedding model.
Load it with a 4,096-token context and full GPU offload, start the local server, and call the
OpenAI-compatible embeddings endpoint:
curl http://127.0.0.1:1234/v1/embeddings \
-H 'Content-Type: application/json' \
-d '{
"model": "nemotron-3-embed-1b-q4",
"input": [
"query: What is retrieval-augmented generation?",
"passage: Retrieval-augmented generation adds retrieved documents to a model prompt."
]
}'
The exact model identifier can differ if LM Studio assigns another load name; check
`http://127.0.0.1:1234/v1/models` when needed.
Retrieval format
Card id model:hf:zenmagnets/Nemotron-3-Embed-1B-Q4_K_M-GGUF · collected 2026-10-02 20:53 UTC · JSON