{"v":1,"id":"model:hf:nvidia/diar_streaming_sortformer_4spk-v2","slug":"model-nvidia-diar-streaming-sortformer-4spk-v2","kind":"model","category":"speech","title":"diar_streaming_sortformer_4spk-v2","summary":"🚨 Announcement: New Model Version Available","source":{"provider":"hf","ref":"nvidia/diar_streaming_sortformer_4spk-v2","url":"https://huggingface.co/nvidia/diar_streaming_sortformer_4spk-v2","rev":"84edd514b8ef68004c10086918cd62f2148cbd59","fetchedAt":"2026-10-02T20:53:12.190Z","etag":"W/\"2bb0-/bmfByDMx6J1LkmwqGg0WBlBn5A\""},"author":{"name":"nvidia","url":"https://huggingface.co/nvidia"},"license":{"spdx":"cc-by-4.0","raw":"cc-by-4.0","open":true},"metrics":{"downloads":34224,"downloadsWeek":15195,"likes":145,"stars":54101,"openIssues":348,"lastRelease":{"tag":"v1.9.4","at":"2026-09-11T05:31:55Z"},"pushedAt":"2026-10-02T13:21:13Z","takenAt":"2026-10-02T20:53:12.190Z"},"tags":["nemo","gguf","speaker-diarization","speaker-recognition","speech","audio","Transformer","FastConformer","Conformer","NEST","pytorch","NeMo","automatic-speech-recognition","dataset:fisher_english","dataset:NIST_SRE_2004-2010","dataset:librispeech","dataset:ami_meeting_corpus","dataset:voxconverse_v0.3","dataset:icsi","dataset:aishell4","dataset:dihard_challenge-3-dev","dataset:NIST_SRE_2000-Disc8_split1","dataset:Alimeeting-train","dataset:DiPCo","model-index"],"pipeline":"automatic-speech-recognition","links":{"github":"ggml-org/whisper.cpp","npm":"nodejs-whisper"},"updatedAt":"2026-09-23T14:47:45.000Z","collectedAt":"2026-10-02T20:53:12.190Z","review":{"numbers":["34,224 downloads on Hugging Face","145 likes","license cc-by-4.0","0.1 GB for diar_streaming_sortformer_4spk-v2.q8_0.gguf","54,101 stars on ggml-org/whisper.cpp","348 open issues and PRs","last release v1.9.4 on 2026-09-11","15,195 npm downloads a week for nodejs-whisper","latest nodejs-whisper@0.3.1"],"log":null},"trust":"unlabeled","health":{"status":"alive","checkedAt":"2026-10-02T20:53:12.190Z","http":200},"description":"# 🚨 Announcement: New Model Version Available\n\nA new version of Sortformer diarization model [Nemotron-3-Diarization](https://huggingface.co/nvidia/nemotron-3-diarization) has been released, supporting 8 speakers and greatly improved accuracy.\n\n# Streaming Sortformer Diarizer 4spk v2\n\nimg {\n display: inline;\n}\n\n[](#model-architecture)\n| [](#model-architecture)\n\nThis model is a streaming version of Sortformer diarizer. [Sortformer](https://arxiv.org/abs/2409.06656)[1] is a novel end-to-end neural model for speaker diarization, trained with unconventional objectives compared to existing end-to-end diarization models.\n\n    \n\n[Streaming Sortformer](https://arxiv.org/abs/2507.18446)[2] employs an Arrival-Order Speaker Cache (AOSC) to store frame-level acoustic embeddings of previously observed speakers.\n\n    \n\n    \n\nSortformer resolves permutation problem in diarization following the arrival-time order of the speech segments from each speaker. \n\nThis speaker diarization model can be used to enable the [NeMo Voice Agent](https://github.com/NVIDIA-NeMo/NeMo/tree/main/examples/voice_agent) to recognize speakers in conversations. See the [NeMo Voice Agent](https://github.com/NVIDIA-NeMo/NeMo/tree/main/examples/voice_agent) and the [YAML configuration](https://github.com/NVIDIA-NeMo/NeMo/blob/316aea20fc2aac4e41bc08123cc77118a9f0b82a/examples/voice_agent/server/server_configs/default.yaml#L25) for more details.\n## Discover more from NVIDIA:\nFor documentation, deployment guides, enterprise-ready APIs, and the latest open models—including Nemotron and other cutting-edge speech, translation, and generative AI—visit the NVIDIA Developer Portal at [developer.nvidia.com](https://developer.nvidia.com/).\nJoin the community to access tools, support, and resources to accelerate your development with NVIDIA’s NeMo, Riva, NIM, and foundation models.\n\n### Explore more from NVIDIA:  \nWhat is [Nemotron](https://www.nv…\n\nSource: https://huggingface.co/nvidia/diar_streaming_sortformer_4spk-v2","install":{"kind":"model","hfId":"nvidia/diar_streaming_sortformer_4spk-v2","gated":false,"format":"gguf","files":[{"name":"diar_streaming_sortformer_4spk-v2.q8_0.gguf","size":147075776,"quant":"Q8_0","sha256":"0679cfeb1ce356d0dea9470b31274f4bfc7eb927497d82005483770666da998a"}],"totalBytes":147075776,"suggestedFile":"diar_streaming_sortformer_4spk-v2.q8_0.gguf","requirements":{"ramGb":1,"diskBytes":147075776,"note":"estimate: suggested file size × 1.15 + 0.5 GB; a real measurement comes with lsh models"},"runWith":["whisper.cpp"],"command":"lsh models install hf:nvidia/diar_streaming_sortformer_4spk-v2"}}