Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones

Maria Teleki; Sai Janjur; Haoran Liu; Oliver Grabner; Ketan Verma; Thomas Docog; Xiangjue Dong; Lingfeng Shi; Cong Wang; Stephanie Birkelbach; Jason Kim; Yin Zhang; Éva Székely; James Caverlee

Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones

Maria Teleki, Sai Janjur, Haoran Liu, Oliver Grabner, Ketan Verma, Thomas Docog, Xiangjue Dong, Lingfeng Shi, Cong Wang, Stephanie Birkelbach, Jason Kim, Yin Zhang, Éva Székely, James Caverlee

TL;DR

This work evaluates proprietary and open-source LLMs across architectures and scales using the DRES evaluation framework and demonstrates that robustness to speech is shaped by specific training objectives.

Abstract

LLMs serve as the backbone in SpeechLLMs, yet their behavior on spontaneous conversational input remains poorly understood. Conversational speech contains pervasive disfluencies -- interjections, edits, and parentheticals -- that are rare in the written corpora used for pre-training. Because gold disfluency removal is a deletion-only task, it serves as a controlled probe to determine whether a model performs faithful structural repair or biased reinterpretation. Using the DRES evaluation framework, we evaluate proprietary and open-source LLMs across architectures and scales. We show that model performance clusters into stable precision-recall regimes reflecting distinct editing policies. Notably, reasoning models systematically over-delete fluent content, revealing a bias toward semantic abstraction over structural fidelity. While fine-tuning achieves SOTA results, it harms generalization. Our findings demonstrate that robustness to speech is shaped by specific training objectives.

Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones

TL;DR

Abstract

Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (6)