The Maivia Gazette

Verified AI news, every morning

Models

Nvidia releases Nemotron 3 Diarization, a free 100M-parameter model that tracks up to eight speakers live

The open-weight model tops VoiceArena's diarization benchmark with a 14.72 percent error rate and cuts errors by 41 percent compared with its predecessor.

An overhead view of a round table with eight seats, each sending out coloured ripples that sometimes overlap.
AI-generated illustration, not event photography. The motion is AI-generated from the still.

Nvidia has released Nemotron 3 Diarization, a model that works out which speaker is talking at each moment of a conversation. It has about 100 million parameters and its weights are freely available. The model can tell apart up to eight speakers, detects when several people talk at the same time, and works on both recordings and live audio. Paired with a speech recognition system such as Nvidia's Parakeet, it can produce transcripts with speaker labels, but the labels are anonymous, such as "speaker_2." On VoiceArena's Diarization-Bench, the model is currently in first place with a diarization error rate of 14.72 percent, ahead of the next-best system at 19.3 percent. The benchmark is strict: overlapping speech counts, and even small misalignments at speaker changes are scored as errors. Compared with its predecessor, Streaming Sortformer, the new model cuts the error rate by 41 percent. Users can set the audio buffer to one of four sizes between 30.4 and 0.32 seconds, and shorter buffers generally lower accuracy. Error rates also rise with more participants, heavy background noise or reverb. Because the model is small and free, developers of meeting transcription, call analytics and captioning tools can run speaker separation themselves rather than rely on a hosted service.

Sources

  1. The DecoderNvidia drops a free 100M-parameter model that identifies up to eight speakers in real timePublished · fetched

Also in this edition