Treat this as a compact engineering check, not a leaderboard. Ten reference/target pairs per language are more useful than a single demo sentence, but still too small to crown a universal winner. The useful reading is where a model is stable or brittle for a specific language and metric.
Fish Audio and Qwen3-TTS often score better on UTMOS even when another model has a higher speaker cosine. Qwen3-TTS Base ICL is included for English, German, Spanish, and Chinese; in this run it is strongest as a naturalness/text-recovery baseline, while dedicated cloning models remain competitive on speaker identity. The right reading is per-language tradeoffs, not a single winner.