SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings
arXiv:2607.01238v1 Announce Type: cross Abstract: Recent advances in speech synthesis have shifted from phoneme representations to direct grapheme modeling. While phonemes address the one-to-many mapp