Store › model › Speech
nemotron-3.5-asr-streaming-0.6b
by nvidia · source Hugging Face · updated 2026-09-10
other708 MB~2 GB RAMsource aliveunlabeled
This model is the multilingual extension of nvidia/nemotron-speech-streaming-en-0.6b, adding language-ID prompt conditioning to support transcription across 40 language-locales from a single model.
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:nvidia/nemotron-3.5-asr-streaming-0.6b and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b
- License: other (custom license: read it at the source before installing) · text
- Requirements: about 2 GB of RAM, 708 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with whisper.cpp.
- Tags:
nemosafetensorsggufnemotron3_5_asrfeature-extractiontransformersspeech-recognitioncache-aware ASRautomatic-speech-recognitionstreaming-asrmultilingualspeechaudioFastConformerRNNTParakeet
Numbers
- 1,240,287 downloads on Hugging Face
- 1,159 likes
- license other
- 0.7 GB for nemotron-3.5-asr-streaming-0.6b.q8_0.gguf
- 18,539 stars on NVIDIA-NeMo/Speech
- 325 open issues and PRs
- last release v3.0.0 on 2026-08-07
Numbers as of 2026-10-02 20:54 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
model.safetensors | 2.4 GB | |
nemotron-3.5-asr-streaming-0.6b.q8_0.gguf | Q8_0 | 708 MB |
From the source README
Nemotron 3.5 ASR
h1, h2, h3, h4, h5, h6 {
color: #76b900; /* NVIDIA green */
font-weight: 700;
}
hr {
border: none;
border-top: 1px solid #e5e7eb;
margin: 2rem 0;
}
/* Improve list spacing */
ul, ol {
margin-top: 0.5rem;
margin-bottom: 0.5rem;
}
/* Badge alignment consistency */
img {
display: inline;
vertical-align: middle;
}
> [!Note]
> This model is the multilingual extension of nvidia/nemotron-speech-streaming-en-0.6b, adding language-ID prompt conditioning to support transcription across 40 language-locales from a single model.
Nemotron 3.5 ASR is a multilingual, streaming Automatic Speech Recognition (ASR) model engineered to deliver high-quality multilingual transcription across both low-latency streaming and high-throughput batch workloads. Developed by NVIDIA, this 600M parameter model transcribes speech into text with native support for punctuation and capitalization, and offers runtime flexibility with configurable chunk sizes, including 80ms, 160ms, 320ms, 560ms, and 1120ms.
By leveraging a state-of-the-art Cache-Aware FastConformer-RNNT architecture, the model eliminates redundant overlapping computations common in traditional "buffered" streaming. This allows it to process only new audio chunks while reusing cached encoder context, significantly improving computational efficiency and minimizing end-to-end delay without sacrificing accuracy.
It was trained on a massive ASR dataset and is engineered to perform across diverse and challenging acoustic conditions.
This model is ready for commercial use.
Release Date
Card id model:hf:nvidia/nemotron-3.5-asr-streaming-0.6b · collected 2026-10-02 20:54 UTC · JSON