Safety

Best-of-N TTS Evaluation is Confounded by ASR Family Alignment

arXiv:2607.08256v1 Announce Type: cross Abstract: Best-of-N (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from N candidates with an automatic speech recognition

DGX agentpaper
safetyarxiv-cs-ai

arXiv:2607.08256v1 Announce Type: cross Abstract: Best-of-N (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from N candidates with an automatic speech recognition (ASR) verifier. We identify an underexplored evaluation confound: a verifier's apparent quality depends strongly on which ASR family judges it. On LibriSpeech-PC test-cleanitep{librispeechpc} with F5-TTSitep{f5tts}, verifier rankings reverse across Whisper, wav2vec~2.0, and HuBERT evaluators, and same-family verifier-evaluator pairs recover 2-3imes more oracle headroom than cross-family pairs despite near-identical representations (linear CKA 0.978) -- a pattern consistent with identity- or lineage-level coupling rather than representational overlap. We propose two extbf{cross-family rank ensembles} (rank-averaging and conjunctive max-rank) that attain the lowest mean WER across three independent evaluators -- 1.61% at N{=}10 (-12% relative to F5-TTS) -- with no measurable degradation under automatic SIM-o/UTMOS metrics; the best single verifier drives WER from 2.06% to 1.72% (-16.5%) under the official F5-TTS evaluator. We recommend cross-evaluator triangulation as default reporting practice.

Source: arXiv cs.AI | 2026-07-10

Loading related sources…