Store › model › Speech
parakeet-tdt-0.6b-v3
by nvidia · source Hugging Face · updated 2026-08-05
cc-by-4.0681 MB~2 GB RAMsource aliveunlabeled
🦜 parakeet-tdt-0.6b-v3: Multilingual Speech-to-Text Model
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:nvidia/parakeet-tdt-0.6b-v3 and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
- License: cc-by-4.0
- Requirements: about 2 GB of RAM, 681 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with whisper.cpp.
- Tags:
transformersnemosafetensorsggufparakeet_tdtfeature-extractionautomatic-speech-recognitionspeechaudioTransducerTransformerTDTFastConformerConformerpytorchNeMo
Numbers
- 569,291 downloads on Hugging Face
- 1,170 likes
- license cc-by-4.0
- 0.7 GB for parakeet-tdt-0.6b-v3.q8_0.gguf
- 18,539 stars on NVIDIA-NeMo/Speech
- 325 open issues and PRs
- last release v3.0.0 on 2026-08-07
Numbers as of 2026-10-02 20:52 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
model.safetensors | 2.3 GB | |
parakeet-tdt-0.6b-v3.q8_0.gguf | Q8_0 | 681 MB |
From the source README
🦜 parakeet-tdt-0.6b-v3: Multilingual Speech-to-Text Model
img {
display: inline;
}
Description:
`parakeet-tdt-0.6b-v3` is a 600-million-parameter multilingual automatic speech recognition (ASR) model designed for high-throughput speech-to-text transcription. It extends the parakeet-tdt-0.6b-v2 model by expanding language support from English to 25 European languages. The model automatically detects the language of the audio and transcribes it without requiring additional prompting. It is part of a series of models that leverage the Granary [1, 2] multilingual corpus as their primary training dataset.
🗣️ Try Demo here: https://huggingface.co/spaces/nvidia/parakeet-tdt-0.6b-v3
Supported Languages:
Bulgarian (bg), Croatian (hr), Czech (cs), Danish (da), Dutch (nl), English (en), Estonian (et), Finnish (fi), French (fr), German (de), Greek (el), Hungarian (hu), Italian (it), Latvian (lv), Lithuanian (lt), Maltese (mt), Polish (pl), Portuguese (pt), Romanian (ro), Slovak (sk), Slovenian (sl), Spanish (es), Swedish (sv), Russian (ru), Ukrainian (uk)
This model is ready for commercial/non-commercial use.
Key Features:
`parakeet-tdt-0.6b-v3`'s key features are built on the foundation of its predecessor, parakeet-tdt-0.6b-v2, and include:
- Automatic punctuation and capitalization
- Accurate word-level and segment-level timestamps
- Long audio transcription, supporting audio up to 24 minutes long with full attention (on A100 80GB) or up to 3 hours with local attention.
- Released under a permissive CC BY 4.0 license
For full details on t…
Source: https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
Card id model:hf:nvidia/parakeet-tdt-0.6b-v3 · collected 2026-10-02 20:52 UTC · JSON