Table of Contents
Fetching ...

Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?

Qixin Deng, Bryan Pardo, Thrasyvoulos N Pappas

TL;DR

The paper addresses whether joint language–audio embeddings encode perceptual timbre semantics by comparing MS-CLAP, LAION-CLAP, and MuQ-MuLan against human timbre descriptors. It employs two experiments: descriptor/instrument-level alignment on an instrument timbre dataset and monotonic trend analysis of DSP-induced timbre changes (EQ and reverb) using SocialFX terms. Results show LAION-CLAP consistently achieves the strongest alignment with human timbre perception across both instrument and effect-based timbres, outperforming the other models. This work highlights the potential for timbre-aware applications and motivates timbre-focused tuning to improve retrieval, manipulation, and generation tasks in multimodal audio-language systems.

Abstract

Understanding and modeling the relationship between language and sound is critical for applications such as music information retrieval,text-guided music generation, and audio captioning. Central to these tasks is the use of joint language-audio embedding spaces, which map textual descriptions and auditory content into a shared embedding space. While multimodal embedding models such as MS-CLAP, LAION-CLAP, and MuQ-MuLan have shown strong performance in aligning language and audio, their correspondence to human perception of timbre, a multifaceted attribute encompassing qualities such as brightness, roughness, and warmth, remains underexplored. In this paper, we evaluate the above three joint language-audio embedding models on their ability to capture perceptual dimensions of timbre. Our findings show that LAION-CLAP consistently provides the most reliable alignment with human-perceived timbre semantics across both instrumental sounds and audio effects.

Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?

TL;DR

The paper addresses whether joint language–audio embeddings encode perceptual timbre semantics by comparing MS-CLAP, LAION-CLAP, and MuQ-MuLan against human timbre descriptors. It employs two experiments: descriptor/instrument-level alignment on an instrument timbre dataset and monotonic trend analysis of DSP-induced timbre changes (EQ and reverb) using SocialFX terms. Results show LAION-CLAP consistently achieves the strongest alignment with human timbre perception across both instrument and effect-based timbres, outperforming the other models. This work highlights the potential for timbre-aware applications and motivates timbre-focused tuning to improve retrieval, manipulation, and generation tasks in multimodal audio-language systems.

Abstract

Understanding and modeling the relationship between language and sound is critical for applications such as music information retrieval,text-guided music generation, and audio captioning. Central to these tasks is the use of joint language-audio embedding spaces, which map textual descriptions and auditory content into a shared embedding space. While multimodal embedding models such as MS-CLAP, LAION-CLAP, and MuQ-MuLan have shown strong performance in aligning language and audio, their correspondence to human perception of timbre, a multifaceted attribute encompassing qualities such as brightness, roughness, and warmth, remains underexplored. In this paper, we evaluate the above three joint language-audio embedding models on their ability to capture perceptual dimensions of timbre. Our findings show that LAION-CLAP consistently provides the most reliable alignment with human-perceived timbre semantics across both instrumental sounds and audio effects.
Paper Structure (6 sections, 1 equation, 3 figures, 2 tables)

This paper contains 6 sections, 1 equation, 3 figures, 2 tables.

Figures (3)

  • Figure 1: similarity vs human ratings per descriptor for MS-CLAP, LAION-CLAP and MuQ-MuLan
  • Figure 2: MS-CLAP, LAION-CLAP and MuQ-MuLan vs human-rated timbre semantic profile for Chinese instruments
  • Figure 3: MS-CLAP, LAION-CLAP and MuQ-MuLan vs human-rated timbre semantic profile for Western instruments