Safety

Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?

arXiv:2510.14249v2 Announce Type: replace-cross Abstract: Understanding and modeling the relationship between language and sound are essential for applications such as music information retrieval, tex

DGX agentpaper
safetyarxiv-cs-ai

arXiv:2510.14249v2 Announce Type: replace-cross Abstract: Understanding and modeling the relationship between language and sound are essential for applications such as music information retrieval, text-guided music generation, and audio captioning. Central to these tasks are joint language-audio embedding spaces, which map textual descriptions and auditory content into a shared representation. Although multimodal embedding models such as MS-CLAP, LAION-CLAP, MuQ-MuLan, and OpenFLAM have shown strong performance in language-audio alignment, their correspondence to human perception of timbre, a multifaceted attribute encompassing qualities such as brightness, roughness, and warmth, remains under-explored. In this paper, we evaluate these joint language-audio embedding models in terms of their ability to capture perceptual timbre semantics. Across two complementary experiments, we find that LAION-CLAP shows relatively strong and consistent alignment with human-perceived timbre semantics across both instrumental sounds and descriptor-conditioned audio effects. At the same time, the overall strength of this alignment remains limited, suggesting that current joint language-audio embeddings capture perceptual timbre semantics only partially. We also observe that, overall, reverb-induced timbre semantics are more consistently encoded than equalization-induced timbre semantics.

Related

Source: arXiv cs.AI | 2026-08-26

Loading related sources…