Towards Fair ASR For Second Language Speakers Using Fairness Prompted Finetuning
Monorama Swain, Bubai Maji, Jagabandhu Mishra, Markus Schedl, Anders Søgaard, Jesper Rindom Jensen
TL;DR
The paper tackles fairness in English ASR for second-language speakers by evaluating Whisper and Seamless-M4T and documenting large WER disparities across 26 accent groups. It introduces fairness-prompted finetuning with lightweight adapters and a multi-objective loss that combines ERM with spectral decoupling, Group-DRO, and invariant risk minimization, achieving a fused objective that reduces group disparities while preserving accuracy. The results show substantial macro-average WER improvements (relative ~58.7% for Whisper and ~58.5% for Seamless-M4T over the large pretrained models; ~9.7% and ~7.8% over standard ERM baselines), and reveal that larger models mitigate but do not eliminate fairness gaps. The approach demonstrates practical potential for more equitable ASR across diverse L2 accents, although a few accents (e.g., Urdu, Indian English in Seamless) remain challenging, pointing to data representation gaps.
Abstract
In this work, we address the challenge of building fair English ASR systems for second-language speakers. Our analysis of widely used ASR models, Whisper and Seamless-M4T, reveals large fluctuations in word error rate (WER) across 26 accent groups, indicating significant fairness gaps. To mitigate this, we propose fairness-prompted finetuning with lightweight adapters, incorporating Spectral Decoupling (SD), Group Distributionally Robust Optimization (Group-DRO), and Invariant Risk Minimization (IRM). Our proposed fusion of traditional empirical risk minimization (ERM) with cross-entropy and fairness-driven objectives (SD, Group DRO, and IRM) enhances fairness across accent groups while maintaining overall recognition accuracy. In terms of macro-averaged word error rate, our approach achieves a relative improvement of 58.7% and 58.5% over the large pretrained Whisper and SeamlessM4T, and 9.7% and 7.8% over them, finetuning with standard empirical risk minimization with cross-entropy loss.
