Blog
Voice cloning benchmarks
July 2, 2026

Voice cloning models, measured across five languages.

We cloned ten FLEURS reference speakers per language, generated ten held-out FLEURS target sentences, then averaged speaker similarity, WER/CER, UTMOS predicted quality, and steady-state speed. The players below show one representative reference and generated output from the same run.

English

FLEURS test/en_us, 10 reference/target pairs

Human ASR floor: 3.5% WER on the original FLEURS target audio.

ModelPrecisionCosineWERUTMOSAudio μRTF med.
OmniVoiceint80.51311.1%2.325.20 s0.17
VoxCPM2bf160.5471.8%2.974.88 s1.11
Chatterbox Multilingualfp160.6421.2%2.225.10 s0.76
Fish Audio S2 Profp160.4893.9%3.474.60 s2.78
Qwen3-TTS Base1.7B bf16 ICL0.4912.4%3.504.95 s1.70

German

FLEURS test/de_de, 10 reference/target pairs

Human ASR floor: 11.4% WER on the original FLEURS target audio.

ModelPrecisionCosineWERUTMOSAudio μRTF med.
OmniVoiceint80.7874.3%2.896.57 s0.22
VoxCPM2bf160.7323.8%2.805.76 s1.09
Chatterbox Multilingualfp160.7785.2%3.326.12 s1.20
Fish Audio S2 Profp160.7187.0%3.345.66 s2.80
Qwen3-TTS Base1.7B bf16 ICL0.72110.8%3.345.79 s1.72

Modern Standard Arabic

FLEURS test/ar_eg, 10 reference/target pairs

Human ASR floor: 32.8% WER on the original FLEURS target audio.

ModelPrecisionCosineWERUTMOSAudio μRTF med.
OmniVoiceint80.72733.2%2.727.04 s0.18
VoxCPM2bf160.67830.2%2.676.26 s1.14
Chatterbox Multilingualfp160.70328.3%2.906.37 s1.10
Fish Audio S2 Profp160.68826.7%3.305.63 s2.83

Spanish

FLEURS test/es_419, 10 reference/target pairs

Human ASR floor: 10.6% WER on the original FLEURS target audio.

ModelPrecisionCosineWERUTMOSAudio μRTF med.
OmniVoiceint80.7589.9%2.866.29 s0.21
VoxCPM2bf160.6415.8%2.624.58 s1.18
Chatterbox Multilingualfp160.7086.9%3.375.08 s1.22
Fish Audio S2 Profp160.6336.8%2.954.48 s2.73
Qwen3-TTS Base1.7B bf16 ICL0.71011.1%3.034.79 s1.76

Chinese

FLEURS test/cmn_hans_cn, 10 reference/target pairs

Human ASR floor: 7.6% CER on the original FLEURS target audio.

ModelPrecisionCosineCERUTMOSAudio μRTF med.
OmniVoiceint80.7094.9%2.836.56 s0.21
VoxCPM2bf160.6954.9%2.915.70 s1.10
Fish Audio S2 Profp160.6346.2%3.554.95 s2.76
Qwen3-TTS Base1.7B bf16 ICL0.7015.7%3.445.24 s1.73

Values are means over ten generated clips per supported language, except RTF, which uses the median to avoid cold-start timing outliers. Higher speaker cosine means the generated clip is closer to the FLEURS reference speaker embedding. Lower WER/CER means the requested text was recovered more cleanly; Chinese is scored with CER, the other languages with WER. Higher UTMOS is a no-reference predicted naturalness score on a 1-5 scale. Lower RTF is faster. A dash means the batch run did not produce reliable timing for that row. These are engineering regression metrics, not a human MOS panel. The same Qwen3-ASR scorer is not equally clean on every language: on the original human FLEURS target clips in this set, its baseline error is 3.5% English WER, 11.4% German WER, 32.8% Arabic WER, 10.6% Spanish WER, and 7.6% Chinese CER. Read WER/CER as a text-recovery stress signal, not as a pure TTS quality score. Precision is listed per row: VoxCPM2’s public full-precision Swift path is bf16, while OmniVoice is shown with the published int8 bundle because the fp16 backbone was not used for this published run. Quantized rows can change both quality and speed, so the table only includes the actual path measured for each row.

Input and output sample rates

SourceRoleRate
FLEURS referencesInput reference16 kHz
OmniVoice / ChatterboxGenerated output24 kHz
Qwen3-TTS Base ICLGenerated output24 kHz
Fish Audio S2 ProGenerated output44.1 kHz
VoxCPM2Generated output48 kHz

The references are 16 kHz, so they only contain observable evidence up to roughly 8 kHz. Higher-rate outputs can sound more open because they preserve or synthesize more high-band breath, sibilance, and room detail, but 44.1 or 48 kHz output does not automatically mean a better clone. Most speaker identity and intelligibility cues live well below that high band, while excessive high-frequency energy can make a voice feel sharp even when WER and speaker cosine look good.

Reference audio and generated clones

English reference

English

FLEURS test/en_us/6306322369645218273.wav

Reference transcript: One can only wonder what the keyboard will become when something newer comes along.

Generated text: While he was working at the hospital Liggins began to investigate premature labor during his spare time.

OmniVoice
int8 clone from the English reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.513
WER μ
11.1%
UTMOS μ
2.32
RTF med.
0.17
VoxCPM2
bf16 clone from the English reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.547
WER μ
1.8%
UTMOS μ
2.97
RTF med.
1.11
Chatterbox Multilingual
fp16 clone from the English reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.642
WER μ
1.2%
UTMOS μ
2.22
RTF med.
0.76
Fish Audio S2 Pro
fp16 clone from the English reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.489
WER μ
3.9%
UTMOS μ
3.47
RTF med.
2.78
Qwen3-TTS Base
1.7B bf16 ICL clone from the English reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.491
WER μ
2.4%
UTMOS μ
3.50
RTF med.
1.70
German reference

German

FLEURS test/de_de/17047810064400454397.wav

Reference transcript: Wenn man eine romanische Sprache spricht, ist es natürlich leichter, Portugiesisch zu erlernen.

Generated text: Die Oberfläche des Mondes besteht aus Gestein und Staub. Die äußere Schicht des Mondes wird als Kruste bezeichnet.

OmniVoice
int8 clone from the German reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.787
WER μ
4.3%
UTMOS μ
2.89
RTF med.
0.22
VoxCPM2
bf16 clone from the German reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.732
WER μ
3.8%
UTMOS μ
2.80
RTF med.
1.09
Chatterbox Multilingual
fp16 clone from the German reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.778
WER μ
5.2%
UTMOS μ
3.32
RTF med.
1.20
Fish Audio S2 Pro
fp16 clone from the German reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.718
WER μ
7.0%
UTMOS μ
3.34
RTF med.
2.80
Qwen3-TTS Base
1.7B bf16 ICL clone from the German reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.721
WER μ
10.8%
UTMOS μ
3.34
RTF med.
1.72
Modern Standard Arabic reference

Modern Standard Arabic

FLEURS test/ar_eg/747984393569561964.wav

Reference transcript: المناطق الكبيرة في الشمال هي مناطق قليلة السكان نوعا ما والبعض غير مأهول تقريبا.

Generated text: قامت فرقة إيروشميث بإلغاء باقي حفلاتهم في جولتهم الموسيقية.

OmniVoice
int8 clone from the Modern Standard Arabic reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.727
WER μ
33.2%
UTMOS μ
2.72
RTF med.
0.18
VoxCPM2
bf16 clone from the Modern Standard Arabic reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.678
WER μ
30.2%
UTMOS μ
2.67
RTF med.
1.14
Chatterbox Multilingual
fp16 clone from the Modern Standard Arabic reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.703
WER μ
28.3%
UTMOS μ
2.90
RTF med.
1.10
Fish Audio S2 Pro
fp16 clone from the Modern Standard Arabic reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.688
WER μ
26.7%
UTMOS μ
3.30
RTF med.
2.83
Spanish reference

Spanish

FLEURS test/es_419/11137049735103408221.wav

Reference transcript: Las escenas se proyectan en las pirámides y todas ellas son iluminadas.

Generated text: La segregación y la recombinación barajan la variación entre uno y otro generación tras generación.

OmniVoice
int8 clone from the Spanish reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.758
WER μ
9.9%
UTMOS μ
2.86
RTF med.
0.21
VoxCPM2
bf16 clone from the Spanish reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.641
WER μ
5.8%
UTMOS μ
2.62
RTF med.
1.18
Chatterbox Multilingual
fp16 clone from the Spanish reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.708
WER μ
6.9%
UTMOS μ
3.37
RTF med.
1.22
Fish Audio S2 Pro
fp16 clone from the Spanish reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.633
WER μ
6.8%
UTMOS μ
2.95
RTF med.
2.73
Qwen3-TTS Base
1.7B bf16 ICL clone from the Spanish reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.710
WER μ
11.1%
UTMOS μ
3.03
RTF med.
1.76
Chinese reference

Chinese

FLEURS test/cmn_hans_cn/6735471425702786619.wav

Reference transcript: 当天,格林尼治时间 (GMT) 约 12 时许,该车辆被拖离事故现场。

Generated text: 植物通过光合作用从阳光中获取养分。它们还能提供荫凉。

OmniVoice
int8 clone from the Chinese reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.709
CER μ
4.9%
UTMOS μ
2.83
RTF med.
0.21
VoxCPM2
bf16 clone from the Chinese reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.695
CER μ
4.9%
UTMOS μ
2.91
RTF med.
1.10
Fish Audio S2 Pro
fp16 clone from the Chinese reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.634
CER μ
6.2%
UTMOS μ
3.55
RTF med.
2.76
Qwen3-TTS Base
1.7B bf16 ICL clone from the Chinese reference. Metrics are ten-clip means; audio is one representative sample.
Cosine μ
0.701
CER μ
5.7%
UTMOS μ
3.44
RTF med.
1.73
Method

Dataset references, not hand-picked demos

References and target texts come from the Google FLEURS test split: English, German, Arabic, Spanish, and Mandarin Chinese. For each language we selected ten reference clips and paired each with a different held-out FLEURS target sentence, so the metrics are less sensitive to one easy or bad sentence. For engines that accept a reference transcript, the exact FLEURS transcript was passed with the audio prompt.

The score shape mirrors the objective side of VoxCPM-style voice cloning evaluation: intelligibility via WER/CER, and cloning via speaker-embedding cosine similarity. The speaker encoder here is Soniqo’s speech embed-speaker --engine mlx, so compare rows inside this table, not against paper SIM percentages directly.

UTMOS is computed with utmos22_strong from SpeechMOS after resampling generated clips to 16 kHz. It gives a no-reference naturalness signal that WER and speaker cosine miss, but it is still a model prediction rather than a human listening study.

Qwen3-TTS Base is measured with the public speech-swift ICL API (Qwen3TTSModel.fromPretrainedWithEncoder + synthesizeWithVoiceCloneICL) using the same FLEURS reference audio and transcript. It is omitted for Arabic because the current Qwen3-TTS language set exposed here does not include Arabic. Chatterbox’s upstream language list includes Chinese, but the current Swift frontend only supports the direct tokenizer path for en, ar, hi, de, es, fr, it, and pt; the Chinese row is intentionally omitted until that frontend lands.

# Public speech-swift CLI example for one generated row.
speech speak "$TEXT" \
  --engine voxcpm2 \
  --voxcpm2-variant bf16 \
  --voxcpm2-ref-audio reference.wav \
  --language arabic \
  --output generated.wav

speech embed-speaker reference.wav --engine mlx --json
speech embed-speaker generated.wav --engine mlx --json
speech transcribe generated.wav --engine qwen3 --model 0.6B --language arabic

Why multiple scores?

A clone can sound like the speaker but say the wrong text, or say the text clearly while missing the speaker. It can also match the text and speaker while carrying audible artifacts. Speaker cosine, WER/CER, and UTMOS catch different failure modes, so all three need to be visible.

Emotion and style attribution

OmniVoice
Broad style hints

Good when you want a cloned speaker with simple delivery guidance, such as a calmer, younger, lower-pitched, or whispered read. This benchmark used a neutral delivery.

Chatterbox Multilingual
Expressiveness strength

Useful when you want the same speaker to sound more restrained or more animated without writing emotion tags into the text. This benchmark kept expressiveness neutral.

VoxCPM2
Voice direction in plain words

Strong fit when you want to describe the target voice or delivery in natural language while still cloning from a reference clip. This benchmark used the reference clip only.

Fish Audio S2 Pro
Acted delivery cues

Best when the script needs explicit moments like laughing, whispering, excitement, or sadness. This benchmark used plain text with no acting cues.

Qwen3-TTS Base
Reference-audio ICL

Uses the FLEURS clip and transcript as in-context conditioning. This is the comparable Qwen3-TTS path for sample-based cloning in the table.

The benchmark intentionally leaves these controls neutral. That keeps speaker similarity tied to the FLEURS reference instead of rewarding a model for adding extra emotion, whispering, shouting, or laughter.

Reading the result

Treat this as a compact engineering check, not a leaderboard. Ten reference/target pairs per language are more useful than a single demo sentence, but still too small to crown a universal winner. The useful reading is where a model is stable or brittle for a specific language and metric.

Fish Audio and Qwen3-TTS often score better on UTMOS even when another model has a higher speaker cosine. Qwen3-TTS Base ICL is included for English, German, Spanish, and Chinese; in this run it is strongest as a naturalness/text-recovery baseline, while dedicated cloning models remain competitive on speaker identity. The right reading is per-language tradeoffs, not a single winner.