The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
Brandon James Carone, Iran R. Roman, Pablo Ripollés
TL;DR
The paper tackles the problem of evaluating deep musical understanding in audio LLMs, arguing that existing benchmarks emphasize surface features over invariant relational representations. It introduces the MUSE Benchmark, a 10-task, open-source suite tested across four SOTA models and a large human baseline ($N=200$), under Standalone and Chain-of-Thought prompting with multiple seeds. Key findings show a wide human–machine gap on abstract and music-theoretic tasks, with some models performing near chance on several tasks; CoT prompting is inconsistently beneficial. The results highlight a need for fundamental advances in training paradigms and architectural design to achieve robust invariant musical representations and reliable deep musical reasoning.
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated capabilities in audio understanding, but current evaluations may obscure fundamental weaknesses in relational reasoning. We introduce the Music Understanding and Structural Evaluation (MUSE) Benchmark, an open-source resource with 10 tasks designed to probe fundamental music perception skills. We evaluate four SOTA models (Gemini Pro and Flash, Qwen2.5-Omni, and Audio-Flamingo 3) against a large human baseline (N=200). Our results reveal a wide variance in SOTA capabilities and a persistent gap with human experts. While Gemini Pro succeeds on basic perception, Qwen and Audio Flamingo 3 perform at or near chance, exposing severe perceptual deficits. Furthermore, we find Chain-of-Thought (CoT) prompting provides inconsistent, often detrimental results. Our work provides a critical tool for evaluating invariant musical representations and driving development of more robust AI systems.
