Store › model › Speech
nemotron-speech-streaming-en-0.6b-gguf
by handy-computer · source Hugging Face · updated 2026-09-15
other453 MB~2 GB RAMsource aliveunlabeled
nemotron-speech-streaming-en-0.6b: transcribe.cpp GGUF
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:handy-computer/nemotron-speech-streaming-en-0.6b-gguf and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/handy-computer/nemotron-speech-streaming-en-0.6b-gguf
- License: other (custom license: read it at the source before installing) · text
- Requirements: about 2 GB of RAM, 453 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with whisper.cpp.
- Tags:
transcribe.cppggufasrspeech-to-textparakeetconformerrnntstreamingcache-awareautomatic-speech-recognitionen
Numbers
- 46,359 downloads on Hugging Face
- 0 likes
- license other
- 0.4 GB for nemotron-speech-streaming-en-0.6b-Q4_K_M.gguf
Numbers as of 2026-10-02 21:00 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
nemotron-speech-streaming-en-0.6b-F16.gguf | F16 | 1.2 GB |
nemotron-speech-streaming-en-0.6b-F32.gguf | F32 | 2.3 GB |
nemotron-speech-streaming-en-0.6b-Q4_K_M.gguf | Q4_K_M | 453 MB |
nemotron-speech-streaming-en-0.6b-Q5_K_M.gguf | Q5_K_M | 514 MB |
nemotron-speech-streaming-en-0.6b-Q6_K.gguf | Q6_K | 573 MB |
nemotron-speech-streaming-en-0.6b-Q8_0.gguf | Q8_0 | 696 MB |
From the source README
nemotron-speech-streaming-en-0.6b: transcribe.cpp GGUF
GGUF conversions of nvidia/nemotron-speech-streaming-en-0.6b for use
with transcribe.cpp.
Ported from upstream commit
ef3bf40,
pinned 2026-05-11.
Validated against the NeMo reference at transcribe.cpp commit
12f1076
on 2026-05-11.
English speech-to-text with punctuation and capitalization. A 0.6B-parameter cache-aware streaming FastConformer encoder with an RNN-T transducer decoder. Runs in both offline and cache-aware streaming modes. The encoder preserves the upstream att_context_size=[70, 13] (1.12s) cache-aware attention mask end-to-end.
Downloads
| Quantization | Download | Size | WER (LibriSpeech test-clean, offline) |
| --- | --- | ---: | ---: |
| F32 | nemotron-speech-streaming-en-0.6b-F32.gguf | 2.47 GB | 2.31% |
| F16 | nemotron-speech-streaming-en-0.6b-F16.gguf | 1.24 GB | 2.31% |
| Q8_0 | nemotron-speech-streaming-en-0.6b-Q8_0.gguf | 730 MB | 2.31% |
| Q6_K | nemotron-speech-streaming-en-0.6b-Q6_K.gguf | 600 MB | 2.29% |
| Q5_K_M | [nemotron-speech-streaming-en-0.6b-Q5_K_M.gguf](https://huggingface.co/handy-c…
Source: https://huggingface.co/handy-computer/nemotron-speech-streaming-en-0.6b-gguf
Card id model:hf:handy-computer/nemotron-speech-streaming-en-0.6b-gguf · collected 2026-10-02 21:00 UTC · JSON