NVIDIA Nemotron 3 Diarization, a 100M-parameter open-weight model, tops VoiceArena's Diarization-Bench leaderboard with a 14.72% Diarization Error Rate. The model identifies who spoke when in live and recorded conversations, supporting up to eight speakers and handling overlapping speech. It builds on NVIDIA's earlier Streaming Sortformer work, extending speaker capacity from four to eight while improving accuracy and throughput. Trained partly on licensed David AI data, it suits speaker-attributed transcription for meetings, calls, and voice agents.
No score is assigned. Sources and their independence are shown in the citation chain below.