Speaker diarization benchmarks: Nemotron 3 and Sortformer
Nemotron 3 MLX and CoreML join earlier diarization systems on the same internal set of 132 recordings across eight languages. Every system receives the audio without a reference speaker count. The Nemotron 3 run covers 12.82 hours of audio with no failed files.
Each cell is DER (diarization error rate), reference-speaker-time weighted with a 0.25 s collar. Lower is better. Real conversation and synthetic rails are different evidence tiers and must not be read across.
Nemotron 3 INT8 exports
The shipped Apple Silicon runtimes use the final NVIDIA Nemotron 3 Diarization checkpoint. Both runs used automatic speaker count and the package's default postprocessing. DER is pooled by reference speaker time inside each recording's evaluation map, including overlap. Exact speaker-count agreement was 85.6% for MLX and 84.8% for CoreML.
| Language | Files | Nemotron 3 MLX | Nemotron 3 CoreML | Sortformer source |
|---|---|---|---|---|
| Arabic | 12 | 16.54% | 18.16% | 22.32% |
| German | 12 | 7.32% | 7.21% | 8.38% |
| English | 48 | 10.31% | 10.24% | 13.81% |
| Spanish | 12 | 12.56% | 13.19% | 16.33% |
| Mandarin | 12 | 10.00% | 9.68% | 9.95% |
| Azerbaijani | 12 | 2.53% | 2.55% | 5.25% |
| Swiss German | 12 | 15.72% | 15.69% | 21.59% |
| Russian | 12 | 4.47% | 4.44% | 5.35% |
| All | 132 | 10.98% | 11.20% | — |
Earlier source-model results by language
Real conversation
Broadcast, telephone and meeting audio. 96 recordings, 12.63 hours.
| System | Arabic | German | English | Spanish | Mandarin |
|---|---|---|---|---|---|
| DER, lower is better | |||||
| Speech Cloud API | 21.02% | 7.62% | 13.21% | 15.92% | 9.63% |
| Sortformer streaming | 22.32% | 8.38% | 13.81% | 16.33% | 9.95% |
| Sortformer offline | 17.92% | 8.13% | 14.34% | 16.43% | 10.52% |
| pyannote Community-1 | 15.57% | 20.22% | 15.22% | 22.78% | 16.19% |
Synthetic rails
Concatenated read speech with turn-taking sampled from distributions measured on the CallHome slices. These detect gross domain failures and nothing more. Do not compare these numbers with the table above.
| System | Azerbaijani | Swiss German | Russian |
|---|---|---|---|
| DER, lower is better | |||
| Speech Cloud API | 4.74% | 19.33% | 8.15% |
| Sortformer streaming | 5.25% | 21.59% | 5.35% |
| Sortformer offline | 5.16% | 28.03% | 8.20% |
| pyannote Community-1 | 35.23% | 45.21% | 35.26% |
Methodology and limits
- One scorer, one item set. Every cell comes from the same harness and the same pinned utterances, so two systems are never compared with two different rulers.
- No speaker count is supplied to any backend.
- Anonymous speaker timelines only. These are diarization results, not named speaker attribution.
- Synthetic rails are not real conversation. The
az,gswandrurails were rebuilt on turn-taking distributions measured from CallHome. Earlier rails used a fixed gap pattern that produced a 100% speaker-change rate, where real conversation repeats a speaker about 16% of the time. That rewarded any diarizer that simply alternates, and it ordered the field wrongly on two languages. - Sortformer caps at four speakers; 22 of the 96 original files exceed that capacity.
Data provenance
The German, Spanish and Mandarin slices come from
talkbank/callhome
at revision 17c8a153 under CC BY-NC-SA 4.0. Reference turns are the
corpus's own timestamps: no VAD, clustering or embedding-derived identity is involved, and overlap
is preserved as published. Because that licence forbids commercial use, these slices are evaluation
evidence and are not training-approved.
Run it yourself
The results use the same benchmark dataset and scorer. See the speaker diarization guide for engine selection, DER scoring and RTTM output on Apple Silicon, or the full benchmark index for the ASR and TTS rows measured the same way.