From the perspective of perceptual speech quality: The robustness of frequency bands to noise
Junyi Fan, Donald S. Williamson
TL;DR
This study investigates how real-world noise affects perceptual speech quality across 32 frequency bands, using a MUSHRA-inspired subjective test to quantify band-level robustness. Stimuli are constructed by filtering speech and noise into bands (100–7500 Hz) and reconstructing signals with target-band and non-target combinations across multiple SNRs, yielding per-band robustness indices $B_{norm}(i,r)$. Results show mid-frequency bands are generally less robust to noise in perceptual quality, while low- and high-frequency bands tend to be more robust, with significant differences across SNRs and noise types; ESTOI aligns more closely with subjective results than PESQ or STOI. The findings inform band-aware strategies for speech enhancement and emphasize the limitations of current objective metrics in predicting perceptual speech quality under realistic noisy conditions, with implications for telecommunications and hearing-impaired applications, and point to the need for better quality metrics and more efficient evaluation methods.
Abstract
Speech quality is one of the main foci of speech-related research, where it is frequently studied with speech intelligibility, another essential measurement. Band-level perceptual speech intelligibility, however, has been studied frequently, whereas speech quality has not been thoroughly analyzed. In this paper, a Multiple Stimuli With Hidden Reference and Anchor (MUSHRA) inspired approach was proposed to study the individual robustness of frequency bands to noise with perceptual speech quality as the measure. Speech signals were filtered into thirty-two frequency bands with compromising real-world noise employed at different signal-to-noise ratios. Robustness to noise indices of individual frequency bands was calculated based on the human-rated perceptual quality scores assigned to the reconstructed noisy speech signals. Trends in the results suggest the mid-frequency region appeared less robust to noise in terms of perceptual speech quality. These findings suggest future research aiming at improving speech quality should pay more attention to the mid-frequency region of the speech signals accordingly.
