Speaker diarization benchmarks: Nemotron 3 and Sortformer

Nemotron 3 MLX and CoreML join earlier diarization systems on the same internal set of 132 recordings across eight languages. Every system receives the audio without a reference speaker count. The Nemotron 3 run covers 12.82 hours of audio with no failed files.

How to read these numbers

Each cell is DER (diarization error rate), reference-speaker-time weighted with a 0.25 s collar. Lower is better. Real conversation and synthetic rails are different evidence tiers and must not be read across.

Nemotron 3 INT8 exports

The shipped Apple Silicon runtimes use the final NVIDIA Nemotron 3 Diarization checkpoint. Both runs used automatic speaker count and the package's default postprocessing. DER is pooled by reference speaker time inside each recording's evaluation map, including overlap. Exact speaker-count agreement was 85.6% for MLX and 84.8% for CoreML.

LanguageFilesNemotron 3 MLXNemotron 3 CoreMLSortformer source
Arabic1216.54%18.16%22.32%
German127.32%7.21%8.38%
English4810.31%10.24%13.81%
Spanish1212.56%13.19%16.33%
Mandarin1210.00%9.68%9.95%
Azerbaijani122.53%2.55%5.25%
Swiss German1215.72%15.69%21.59%
Russian124.47%4.44%5.35%
All13210.98%11.20%—

Earlier source-model results by language

Real conversation

Broadcast, telephone and meeting audio. 96 recordings, 12.63 hours.

SystemArabicGermanEnglishSpanishMandarin
DER, lower is better
Speech Cloud API21.02%7.62%13.21%15.92%9.63%
Sortformer streaming22.32%8.38%13.81%16.33%9.95%
Sortformer offline17.92%8.13%14.34%16.43%10.52%
pyannote Community-115.57%20.22%15.22%22.78%16.19%

Synthetic rails

Concatenated read speech with turn-taking sampled from distributions measured on the CallHome slices. These detect gross domain failures and nothing more. Do not compare these numbers with the table above.

SystemAzerbaijaniSwiss GermanRussian
DER, lower is better
Speech Cloud API4.74%19.33%8.15%
Sortformer streaming5.25%21.59%5.35%
Sortformer offline5.16%28.03%8.20%
pyannote Community-135.23%45.21%35.26%

Methodology and limits

Data provenance

The German, Spanish and Mandarin slices come from talkbank/callhome at revision 17c8a153 under CC BY-NC-SA 4.0. Reference turns are the corpus's own timestamps: no VAD, clustering or embedding-derived identity is involved, and overlap is preserved as published. Because that licence forbids commercial use, these slices are evaluation evidence and are not training-approved.

Run it yourself

The results use the same benchmark dataset and scorer. See the speaker diarization guide for engine selection, DER scoring and RTTM output on Apple Silicon, or the full benchmark index for the ASR and TTS rows measured the same way.