Research
Your Voice Cloning System is Secretly a Voice Anonymizer
arXiv:2608.27360v1 Announce Type: new Abstract: Speaker anonymization suppresses speaker-identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTS
arXiv:2608.27360v1 Announce Type: new Abstract: Speaker anonymization suppresses speaker-identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speech, for speaker anonymization without retraining. Our key insight is that XTTSv2's voice cloning capabilities preserve prosodic structure independently of speaker identity, enabling voice conversion by conditioning on a pseudo-speaker. We introduce an iterative refinement strategy that balances privacy and utility by maximizing a harmonic mean of speaker dissimilarity and intelligibility. Evaluated on seven European languages across CommonVoice and Multilingual LibriSpeech, our system achieves near-optimal privacy (EER approx 0.49), competitive intelligibility, and substantially better speech quality than dedicated anonymization baselines, while requiring no language-specific training. We release the code here: https://github.com/rm00cr/coqui-tts.
Related
- You Are What You Say: Exploiting Linguistic Content for VoicePrivacy Attacks
- Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models
- Content Anonymization for Privacy in Long-form Audio
Source: arXiv cs.CL | 2026-08-28