Model Releases
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
arXiv:2607.27109v2 Announce Type: cross Abstract: With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-gra
arXiv:2607.27109v2 Announce Type: cross Abstract: With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a extbf{M}assive extbf{M}ulti-dimensional benchmark for extbf{A}udio extbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the reference label. We evaluate representative open-source and proprietary AudioLLMs. Results show clear differences across evaluation dimensions, information coverage, and description reliability. We will release the MMAC benchmark and evaluation code.
Related
- MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
- ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
- AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs
- Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
Source: arXiv cs.AI | 2026-07-31