Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models
DGX agentarXiv:2604.15383v1 Announce Type: cross Abstract: Large audio-language models (LALMs) generalize across speech, sound, and music, but unified decoders can exhibit a temporal smoothing bias: transient