Adaptive Per-Channel Energy Normalization Front-end for Robust Audio Signal Processing
Hanyu Meng, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Qiquan Zhang, Haizhou Li
TL;DR
This paper tackles the inflexibility of fixed learnable audio front-ends by introducing LEAF-APCEN, an adaptive front-end that uses a neural controller to modulate a simplified PCEN (SimpPCEN) within the LEAF framework. The controller reads current and past subband energies to produce time-varying parameters that adjust dynamic range compression in a per-subband manner, yielding input-dependent representations. Empirical results across environmental sound, music, emotion, and speaker identification tasks show that LEAF-APCEN achieves higher accuracy and substantially better robustness under complex acoustic conditions, with faster convergence and more discriminative feature representations. The work demonstrates the value of neural adaptability in audio front-ends and points toward future multi-channel extensions that integrate temporal, spectral, and spatial adaptation for even more robust performance.
Abstract
In audio signal processing, learnable front-ends have shown strong performance across diverse tasks by optimizing task-specific representation. However, their parameters remain fixed once trained, lacking flexibility during inference and limiting robustness under dynamic complex acoustic environments. In this paper, we introduce a novel adaptive paradigm for audio front-ends that replaces static parameterization with a closed-loop neural controller. Specifically, we simplify the learnable front-end LEAF architecture and integrate a neural controller for adaptive representation via dynamically tuning Per-Channel Energy Normalization. The neural controller leverages both the current and the buffered past subband energies to enable input-dependent adaptation during inference. Experimental results on multiple audio classification tasks demonstrate that the proposed adaptive front-end consistently outperforms prior fixed and learnable front-ends under both clean and complex acoustic conditions. These results highlight neural adaptability as a promising direction for the next generation of audio front-ends.
