What do MLLMs hear? Examining reasoning with text and sound components in Multimodal Large Language Models

Enis Berk Çoban; Michael I. Mandel; Johanna Devaney

What do MLLMs hear? Examining reasoning with text and sound components in Multimodal Large Language Models

Enis Berk Çoban, Michael I. Mandel, Johanna Devaney

TL;DR

This work interrogates whether reasoning capabilities of LLMs can be harnessed by multimodal LLMs for audio-based classification and relational reasoning. Through Experiment 1, the authors evaluate in-context learning with the LTU audio MLLM on the EDANSA dataset, finding that caption-based reasoning benefits from full fine-tuning but prompting alone offers limited gains. Experiment 2 probes semantic concept representations using synonyms and hypernyms, revealing robust text-only reasoning for synonyms but substantial cross-modal gaps for hypernym-based hierarchical relationships when audio is involved. Overall, the study demonstrates significant cross-modal alignment limitations in current audio MLLMs and highlights the need for finer-grained audio-text grounding to realize true co-reasoning across modalities.

Abstract

Large Language Models (LLMs) have demonstrated remarkable reasoning capabilities, notably in connecting ideas and adhering to logical rules to solve problems. These models have evolved to accommodate various data modalities, including sound and images, known as multimodal LLMs (MLLMs), which are capable of describing images or sound recordings. Previous work has demonstrated that when the LLM component in MLLMs is frozen, the audio or visual encoder serves to caption the sound or image input facilitating text-based reasoning with the LLM component. We are interested in using the LLM's reasoning capabilities in order to facilitate classification. In this paper, we demonstrate through a captioning/classification experiment that an audio MLLM cannot fully leverage its LLM's text-based reasoning when generating audio captions. We also consider how this may be due to MLLMs separately representing auditory and textual information such that it severs the reasoning pathway from the LLM to the audio encoder.

What do MLLMs hear? Examining reasoning with text and sound components in Multimodal Large Language Models

TL;DR

Abstract

Paper Structure (19 sections, 2 figures, 3 tables)

This paper contains 19 sections, 2 figures, 3 tables.

Introduction
Reasoning in multimodal large language models
Visual reasoning in MLLMs
Audio reasoning in MLLMs
Experiment 1: In-context audio classification
Methodology
In-Context Learning with Grouse Call Descriptions
Results
Discussion
Experiment 2: Examining concept representations in an audio MLLM
Methodology
Results
Discussion
Limitations
Conclusions and Future Work
...and 4 more sections

Figures (2)

Figure 1: Generic audio MLLM architecture, specific components may vary with specific models. The snowflake represents components that are typically frozen and the flame represents those that are typically trained (or can be in the case of the '/').
Figure 2: Results of experiment on concept representations in an audio MLLM. Subplot (a) shows the results for the similarity (synonyms and unrelated terms) category and subplot (b) for the hierarchy (hypernym) category. Each subplot shows violin plots of the yes rate for four different conditions: text-only prompting, text prompting with a silent audio file, text prompting with an audio file from AudioSet, and text prompting with an audio file from EDANSA. The prompts (P1–P4) are defined in in Table \ref{['exp2prompts']}.

What do MLLMs hear? Examining reasoning with text and sound components in Multimodal Large Language Models

TL;DR

Abstract

What do MLLMs hear? Examining reasoning with text and sound components in Multimodal Large Language Models

Authors

TL;DR

Abstract

Table of Contents

Figures (2)