LogiShell store Open app

Store › model › Speech

multitalker-parakeet-streaming-0.6b-v1-gguf

by handy-computer · source Hugging Face · updated 2026-09-15

other589 MB~2 GB RAMsource aliveunlabeled

multitalker-parakeet-streaming-0.6b-v1: transcribe.cpp GGUF

Add to LogiShell Open in the web IDE

The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:handy-computer/multitalker-parakeet-streaming-0.6b-v1-gguf and its progress lives in the Resource Center.

Source and license

Numbers

Numbers as of 2026-10-02 21:00 UTC, from the source API.

Summary

Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.

Reviews

No reviews yet. Reviews are written inside LogiShell: open this card in the app.

Files

filequantsize
bundle/multitalker-parakeet-streaming-0.6b-v1-F16.ggufF161.4 GB
bundle/multitalker-parakeet-streaming-0.6b-v1-F32.ggufF322.8 GB
bundle/multitalker-parakeet-streaming-0.6b-v1-Q4_K_M.ggufQ4_K_M589 MB
bundle/multitalker-parakeet-streaming-0.6b-v1-Q5_K_M.ggufQ5_K_M650 MB
bundle/multitalker-parakeet-streaming-0.6b-v1-Q6_K.ggufQ6_K709 MB
bundle/multitalker-parakeet-streaming-0.6b-v1-Q8_0.ggufQ8_0833 MB

From the source README

multitalker-parakeet-streaming-0.6b-v1: transcribe.cpp GGUF

GGUF conversions of nvidia/multitalker-parakeet-streaming-0.6b-v1 for use
with transcribe.cpp.

Ported from upstream commit
8749fc7,
pinned 2026-07-12.
Validated against the NeMo reference at transcribe.cpp commit
3083021
on 2026-08-03.

Offline and cache-aware streaming English speech-to-text with punctuation and capitalization. A 0.6B-parameter cache-aware streaming FastConformer encoder with an RNN-T transducer decoder, fine-tuned from nvidia/nemotron-speech-streaming-en-0.6b. Plain GGUFs run the single_speaker_mode ASR path, while bundle GGUFs under `bundle/` embed nvidia/diar_streaming_sortformer_4spk-v2.1 and, with `--diarize`, transcribe up to four overlapping speakers into a speaker-tagged transcript. The encoder preserves the upstream att_context_size=[70, 13] (1.12s) cache-aware attention mask; all four latency lookahead settings are selectable.

Downloads

| Quantization | Download | Size | WER (LibriSpeech test-clean, offline) |
| --- | --- | ---: | ---: |
| F32 | bundle/multitalker-parakeet-streaming-0.6b-v1-F32.gguf | 2.96 GB | 2.19% |
| F16 | bundle/multitalker-parakeet-streaming-0.6b-v1-F16.gguf | 1.48 GB | 2.19% |
| Q8_0 | [bundle/multitalker-parakeet-streaming-0.6b-v1-Q8_0.gguf](https://huggingface.co/handy-computer/multit…

Source: https://huggingface.co/handy-computer/multitalker-parakeet-streaming-0.6b-v1-gguf

Card id model:hf:handy-computer/multitalker-parakeet-streaming-0.6b-v1-gguf · collected 2026-10-02 21:00 UTC · JSON