How Much Can AI Hear? MMAC Tests 5,638 Audio Clips Across 15 Dimensions

MMAC spans 5,638 clips from more than 20 sources and separates whether a caption covers relevant information from whether its claims match the reference.

Beyond a short audio caption

MMAC, submitted July 29, evaluates fine-grained free-form descriptions from AudioLLMs. Existing metrics often emphasize sentence quality or one task score, making it hard to separate omitted information from incorrect claims.

Data and evaluation

The benchmark contains 5,638 clips from more than 20 sources, covering six capability categories and fifteen evaluation dimensions. For each generated caption, it separately checks whether relevant information is mentioned and whether the mentioned content agrees with a reference label.

The authors evaluated representative open and proprietary AudioLLMs and report clear differences in dimensional coverage and description reliability. They plan to release the benchmark and evaluation code.

Why it matters and limitations

For accessibility captions, safety monitoring, and media search, knowing which event, source, or spatial cue was missed matters more than receiving one fluent sentence. Developers should inspect dimension-level omissions and unsupported assertions alongside an overall score.

The abstract provides no detailed per-model results, so the current public information does not support a model ranking. Source-language and environment bias, along with imperfect human reference labels, require examination. The paper is not yet peer reviewed.

Primary source