{"v":1,"id":"model:hf:zenmagnets/Nemotron-3-Embed-1B-Q4_K_M-GGUF","slug":"model-zenmagnets-nemotron-3-embed-1b-q4-k-m-gguf","kind":"model","category":"embedding","title":"Nemotron-3-Embed-1B-Q4_K_M-GGUF","summary":"An independently converted and quantized GGUF of NVIDIA's nvidia/Nemotron-3-Embed-1B-BF16, prepared for local embedding inference in LM Studio and llama.cpp-compatible runtimes.","source":{"provider":"hf","ref":"zenmagnets/Nemotron-3-Embed-1B-Q4_K_M-GGUF","url":"https://huggingface.co/zenmagnets/Nemotron-3-Embed-1B-Q4_K_M-GGUF","rev":"06df1fde6f7009c91f6cc3cd520081921929a678","fetchedAt":"2026-10-02T20:53:29.167Z","etag":"W/\"a28-kDRdBHGmvBn3aXp+tWYriTxzbzg\""},"author":{"name":"zenmagnets","url":"https://huggingface.co/zenmagnets"},"license":{"spdx":"openmdw-1.1","raw":"openmdw-1.1","open":true},"metrics":{"downloads":102925,"likes":5,"stars":18539,"openIssues":325,"lastRelease":{"tag":"v3.0.0","at":"2026-08-07T00:13:21Z"},"pushedAt":"2026-10-02T16:11:38Z","takenAt":"2026-10-02T20:53:29.167Z"},"tags":["gguf","Q4_K_M","text-embeddings","feature-extraction","retrieval","semantic-search","rag","lm-studio","llama-cpp","sentence-similarity","multilingual","en","ar","as","bn","bg","zh","da","nl","fi","fr","de","hi","id","it","ja","ko","ms","mr","ne","no","fa","pt","ro","ru","es","sw","sv","ta","te"],"pipeline":"sentence-similarity","links":{"github":"NVIDIA-NeMo/Speech"},"updatedAt":"2026-07-17T05:15:42.000Z","collectedAt":"2026-10-02T20:53:29.167Z","review":{"numbers":["102,925 downloads on Hugging Face","5 likes","license openmdw-1.1","0.7 GB for nemotron-3-embed-1b-q4_k_m.gguf","18,539 stars on NVIDIA-NeMo/Speech","325 open issues and PRs","last release v3.0.0 on 2026-08-07"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:53:29.167Z","http":200},"description":"# Nemotron-3-Embed-1B Q4_K_M GGUF\n\nAn independently converted and quantized GGUF of NVIDIA's\n[`nvidia/Nemotron-3-Embed-1B-BF16`](https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16),\nprepared for local embedding inference in LM Studio and llama.cpp-compatible runtimes.\n\nThis repository is not an official NVIDIA release and is not affiliated with or endorsed by NVIDIA.\n\n## File\n\n| File | Quantization | Size | SHA-256 |\n|---|---:|---:|---|\n| `nemotron-3-embed-1b-q4_k_m.gguf` | Q4_K_M | 749,352,096 bytes (714.6 MiB) | `9a74166f51dbc280073748fa199bea49283bd21f7f9280f2dec2b4d975ddfd1d` |\n\nThe model produces 2,048-dimensional, L2-normalized embeddings. Its GGUF metadata declares a\n262,144-token maximum context. The release was functionally tested at a 4,096-token context; very\nlarge contexts were not validated and require substantially more memory.\n\n## Use with LM Studio\n\nDownload the GGUF from **Files and versions**, then drag it into LM Studio or place it in LM Studio's\nmodels directory. LM Studio should classify it as an embedding model.\n\nLoad it with a 4,096-token context and full GPU offload, start the local server, and call the\nOpenAI-compatible embeddings endpoint:\n\n```bash\ncurl http://127.0.0.1:1234/v1/embeddings \\\n  -H 'Content-Type: application/json' \\\n  -d '{\n    \"model\": \"nemotron-3-embed-1b-q4\",\n    \"input\": [\n      \"query: What is retrieval-augmented generation?\",\n      \"passage: Retrieval-augmented generation adds retrieved documents to a model prompt.\"\n    ]\n  }'\n```\n\nThe exact model identifier can differ if LM Studio assigns another load name; check\n`http://127.0.0.1:1234/v1/models` when needed.\n\n## Retrieval format\n\nUse the prefixes specified by NVIDIA:\n\n- Queries: `query: `\n- Documents: `passage: `\n\nEmbeddings are normalized, so cosine similarity and dot product give equivalent rankings (within\nnormal floating-point tolerance).\n\n## Conversion provenance\n\n- Upstream repository…\n\nSource: https://huggingface.co/zenmagnets/Nemotron-3-Embed-1B-Q4_K_M-GGUF","install":{"kind":"model","hfId":"zenmagnets/Nemotron-3-Embed-1B-Q4_K_M-GGUF","gated":false,"format":"gguf","files":[{"name":"nemotron-3-embed-1b-q4_k_m.gguf","size":749352096,"quant":"Q4_K_M","sha256":"9a74166f51dbc280073748fa199bea49283bd21f7f9280f2dec2b4d975ddfd1d"}],"totalBytes":749352096,"suggestedFile":"nemotron-3-embed-1b-q4_k_m.gguf","requirements":{"ramGb":2,"diskBytes":749352096,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["llama.cpp"],"command":"lsh models install hf:zenmagnets/Nemotron-3-Embed-1B-Q4_K_M-GGUF"}}