Store › model › Speech
diar_streaming_sortformer_4spk-v2
by nvidia · source Hugging Face · updated 2026-09-23
cc-by-4.0140 MB~1 GB RAMsource aliveunlabeled
🚨 Announcement: New Model Version Available
Add to LogiShell Open in the web IDE
The button opens LogiShell with this card; nothing installs from a link by itself. Inside the app the install goes through lsh models install hf:nvidia/diar_streaming_sortformer_4spk-v2 and its progress lives in the Resource Center.
Source and license
- Source: https://huggingface.co/nvidia/diar_streaming_sortformer_4spk-v2
- License: cc-by-4.0
- Requirements: about 1 GB of RAM, 140 MB on disk (estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models). Runs with whisper.cpp.
- Tags:
nemoggufspeaker-diarizationspeaker-recognitionspeechaudioTransformerFastConformerConformerNESTpytorchNeMoautomatic-speech-recognitiondataset:fisher_englishdataset:NIST_SRE_2004-2010dataset:librispeech
Numbers
- 34,224 downloads on Hugging Face
- 145 likes
- license cc-by-4.0
- 0.1 GB for diar_streaming_sortformer_4spk-v2.q8_0.gguf
- 54,101 stars on ggml-org/whisper.cpp
- 348 open issues and PRs
- last release v1.9.4 on 2026-09-11
- 15,195 npm downloads a week for nodejs-whisper
- latest nodejs-whisper@0.3.1
Numbers as of 2026-10-02 20:53 UTC, from the source API.
Summary
Summary not ready yet: the numbers are here, the text is not. It is written by the collector through the LogiShell model facade when a provider key is present.
Reviews
No reviews yet. Reviews are written inside LogiShell: open this card in the app.
Files
| file | quant | size |
|---|---|---|
diar_streaming_sortformer_4spk-v2.q8_0.gguf | Q8_0 | 140 MB |
From the source README
🚨 Announcement: New Model Version Available
A new version of Sortformer diarization model Nemotron-3-Diarization has been released, supporting 8 speakers and greatly improved accuracy.
Streaming Sortformer Diarizer 4spk v2
img {
display: inline;
}
This model is a streaming version of Sortformer diarizer. Sortformer[1] is a novel end-to-end neural model for speaker diarization, trained with unconventional objectives compared to existing end-to-end diarization models.
Streaming Sortformer[2] employs an Arrival-Order Speaker Cache (AOSC) to store frame-level acoustic embeddings of previously observed speakers.
Sortformer resolves permutation problem in diarization following the arrival-time order of the speech segments from each speaker.
This speaker diarization model can be used to enable the NeMo Voice Agent to recognize speakers in conversations. See the NeMo Voice Agent and the YAML configuration for more details.
## Discover more from NVIDIA:
For documentation, deployment guides, enterprise-ready APIs, and the latest open models—including Nemotron and other cutting-edge speech, translation, and generative AI—visit the NVIDIA Developer Portal at developer.nvidia.com.
Join the community to access tools, support, and resources to accelerate your development with NVIDIA’s NeMo, Riva, NIM, and foundation models.
### Explore more from NVIDIA:
What is [Nemotron](https://www.nv…
Source: https://huggingface.co/nvidia/diar_streaming_sortformer_4spk-v2
Card id model:hf:nvidia/diar_streaming_sortformer_4spk-v2 · collected 2026-10-02 20:53 UTC · JSON