Table of Contents
Fetching ...

The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS

Brandon James Carone, Iran R. Roman, Pablo Ripollés

TL;DR

The paper tackles the problem of evaluating deep musical understanding in audio LLMs, arguing that existing benchmarks emphasize surface features over invariant relational representations. It introduces the MUSE Benchmark, a 10-task, open-source suite tested across four SOTA models and a large human baseline ($N=200$), under Standalone and Chain-of-Thought prompting with multiple seeds. Key findings show a wide human–machine gap on abstract and music-theoretic tasks, with some models performing near chance on several tasks; CoT prompting is inconsistently beneficial. The results highlight a need for fundamental advances in training paradigms and architectural design to achieve robust invariant musical representations and reliable deep musical reasoning.

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated capabilities in audio understanding, but current evaluations may obscure fundamental weaknesses in relational reasoning. We introduce the Music Understanding and Structural Evaluation (MUSE) Benchmark, an open-source resource with 10 tasks designed to probe fundamental music perception skills. We evaluate four SOTA models (Gemini Pro and Flash, Qwen2.5-Omni, and Audio-Flamingo 3) against a large human baseline (N=200). Our results reveal a wide variance in SOTA capabilities and a persistent gap with human experts. While Gemini Pro succeeds on basic perception, Qwen and Audio Flamingo 3 perform at or near chance, exposing severe perceptual deficits. Furthermore, we find Chain-of-Thought (CoT) prompting provides inconsistent, often detrimental results. Our work provides a critical tool for evaluating invariant musical representations and driving development of more robust AI systems.

The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS

TL;DR

The paper tackles the problem of evaluating deep musical understanding in audio LLMs, arguing that existing benchmarks emphasize surface features over invariant relational representations. It introduces the MUSE Benchmark, a 10-task, open-source suite tested across four SOTA models and a large human baseline (), under Standalone and Chain-of-Thought prompting with multiple seeds. Key findings show a wide human–machine gap on abstract and music-theoretic tasks, with some models performing near chance on several tasks; CoT prompting is inconsistently beneficial. The results highlight a need for fundamental advances in training paradigms and architectural design to achieve robust invariant musical representations and reliable deep musical reasoning.

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated capabilities in audio understanding, but current evaluations may obscure fundamental weaknesses in relational reasoning. We introduce the Music Understanding and Structural Evaluation (MUSE) Benchmark, an open-source resource with 10 tasks designed to probe fundamental music perception skills. We evaluate four SOTA models (Gemini Pro and Flash, Qwen2.5-Omni, and Audio-Flamingo 3) against a large human baseline (N=200). Our results reveal a wide variance in SOTA capabilities and a persistent gap with human experts. While Gemini Pro succeeds on basic perception, Qwen and Audio Flamingo 3 perform at or near chance, exposing severe perceptual deficits. Furthermore, we find Chain-of-Thought (CoT) prompting provides inconsistent, often detrimental results. Our work provides a critical tool for evaluating invariant musical representations and driving development of more robust AI systems.
Paper Structure (14 sections, 2 figures, 2 tables)

This paper contains 14 sections, 2 figures, 2 tables.

Figures (2)

  • Figure 1: SOTA model comparison on the MUSE benchmark. Models shown with solid lines. Humans shown with dashed and dotted lines.
  • Figure 2: Relationship between number of shots provided and accuracy across Gemini Pro and Gemini Flash (black). The relationship between human accuracy and musical training is also shown (blue). Points represent the estimated effect size (log-odds ratio) from a GLM for the models, and a GLMER for the humans, for each task. Error bars indicate the 95% confidence interval. Positive estimates mean greater shots or training correspond to higher accuracy. The shape of each point indicates the statistical significance of the effect.