{"v":1,"id":"model:hf:handy-computer/nemotron-speech-streaming-en-0.6b-gguf","slug":"model-handy-computer-nemotron-speech-streaming-en-0-6b-gguf","kind":"model","category":"speech","title":"nemotron-speech-streaming-en-0.6b-gguf","summary":"nemotron-speech-streaming-en-0.6b: transcribe.cpp GGUF","source":{"provider":"hf","ref":"handy-computer/nemotron-speech-streaming-en-0.6b-gguf","url":"https://huggingface.co/handy-computer/nemotron-speech-streaming-en-0.6b-gguf","rev":"9789e0ebf77277911272f0d9a35e1646b5aa6004","fetchedAt":"2026-10-02T21:00:06.538Z","etag":"W/\"d82-+12c6aKuiHzXY7+rcDjc3CQQSUA\""},"author":{"name":"handy-computer","url":"https://huggingface.co/handy-computer"},"license":{"spdx":null,"raw":"other","url":"https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/","open":null,"note":"custom license: read it at the source before installing"},"metrics":{"downloads":46359,"likes":0,"takenAt":"2026-10-02T21:00:06.538Z"},"tags":["transcribe.cpp","gguf","asr","speech-to-text","parakeet","conformer","rnnt","streaming","cache-aware","automatic-speech-recognition","en"],"pipeline":"automatic-speech-recognition","links":{"github":"NVIDIA-NeMo/NeMo"},"updatedAt":"2026-09-15T07:05:40.000Z","collectedAt":"2026-10-02T21:00:06.538Z","review":{"numbers":["46,359 downloads on Hugging Face","0 likes","license other","0.4 GB for nemotron-speech-streaming-en-0.6b-Q4_K_M.gguf"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T21:00:06.538Z","http":200},"description":"# nemotron-speech-streaming-en-0.6b: transcribe.cpp GGUF\n\nGGUF conversions of [nvidia/nemotron-speech-streaming-en-0.6b](https://huggingface.co/nvidia/nemotron-speech-streaming-en-0.6b) for use\nwith [transcribe.cpp](https://github.com/handy-computer/transcribe.cpp).\n\nPorted from upstream commit\n[ef3bf40](https://huggingface.co/nvidia/nemotron-speech-streaming-en-0.6b/commit/ef3bf40),\npinned 2026-05-11.\nValidated against the NeMo reference at transcribe.cpp commit\n[12f1076](https://github.com/handy-computer/transcribe.cpp/tree/12f1076)\non 2026-05-11.\n\nEnglish speech-to-text with punctuation and capitalization. A 0.6B-parameter cache-aware streaming FastConformer encoder with an RNN-T transducer decoder. Runs in both offline and cache-aware streaming modes. The encoder preserves the upstream att_context_size=[70, 13] (1.12s) cache-aware attention mask end-to-end.\n\n## Downloads\n\n| Quantization | Download | Size | WER (LibriSpeech test-clean, offline) |\n| --- | --- | ---: | ---: |\n| F32 | [nemotron-speech-streaming-en-0.6b-F32.gguf](https://huggingface.co/handy-computer/nemotron-speech-streaming-en-0.6b-gguf/resolve/main/nemotron-speech-streaming-en-0.6b-F32.gguf) | 2.47 GB | 2.31% |\n| F16 | [nemotron-speech-streaming-en-0.6b-F16.gguf](https://huggingface.co/handy-computer/nemotron-speech-streaming-en-0.6b-gguf/resolve/main/nemotron-speech-streaming-en-0.6b-F16.gguf) | 1.24 GB | 2.31% |\n| Q8_0 | [nemotron-speech-streaming-en-0.6b-Q8_0.gguf](https://huggingface.co/handy-computer/nemotron-speech-streaming-en-0.6b-gguf/resolve/main/nemotron-speech-streaming-en-0.6b-Q8_0.gguf) | 730 MB | 2.31% |\n| Q6_K | [nemotron-speech-streaming-en-0.6b-Q6_K.gguf](https://huggingface.co/handy-computer/nemotron-speech-streaming-en-0.6b-gguf/resolve/main/nemotron-speech-streaming-en-0.6b-Q6_K.gguf) | 600 MB | 2.29% |\n| Q5_K_M | [nemotron-speech-streaming-en-0.6b-Q5_K_M.gguf](https://huggingface.co/handy-c…\n\nSource: https://huggingface.co/handy-computer/nemotron-speech-streaming-en-0.6b-gguf","install":{"kind":"model","hfId":"handy-computer/nemotron-speech-streaming-en-0.6b-gguf","gated":false,"format":"gguf","files":[{"name":"nemotron-speech-streaming-en-0.6b-F16.gguf","size":1237652608,"quant":"F16","sha256":"dc8c4bd7dce6e1805b4e00bef391c3870aca4fa99a563ec02a6f5e0f15d2f489"},{"name":"nemotron-speech-streaming-en-0.6b-F32.gguf","size":2472386176,"quant":"F32","sha256":"53bdc1d4d21e419d4da513d2df31ce88784e25cc908841386490658854b13b41"},{"name":"nemotron-speech-streaming-en-0.6b-Q4_K_M.gguf","size":475436032,"quant":"Q4_K_M","sha256":"dc959ca31499b114e395c44eb4f0778968f20e5cfb03305a08a39925b2da8e1e"},{"name":"nemotron-speech-streaming-en-0.6b-Q5_K_M.gguf","size":538989568,"quant":"Q5_K_M","sha256":"d6b8b36f9aca6a751779d064a29f6583f3e82c0c15646904383c8436a25c1489"},{"name":"nemotron-speech-streaming-en-0.6b-Q6_K.gguf","size":600420352,"quant":"Q6_K","sha256":"b31fcc5e9f3b00cb33c10ff0f05d7e71265f160eeed6b497a65cd296dddd21f3"},{"name":"nemotron-speech-streaming-en-0.6b-Q8_0.gguf","size":729650176,"quant":"Q8_0","sha256":"90d8c89714cd31efc88be62a40c6b2bea57e0cc2063af1ffe2c28f1a228ca110"}],"totalBytes":6054534912,"suggestedFile":"nemotron-speech-streaming-en-0.6b-Q4_K_M.gguf","requirements":{"ramGb":2,"diskBytes":475436032,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["whisper.cpp"],"command":"lsh models install hf:handy-computer/nemotron-speech-streaming-en-0.6b-gguf"}}