PhoenixCodec: Taming Neural Speech Coding for Extreme Low-Resource Scenarios
Zixiang Wan, Haoran Zhao, Guochang Zhang, Runqiang Han, Jianqiang Wei, Yuexian Zou
TL;DR
PhoenixCodec tackles extreme low-resource neural speech coding by unifying an asymmetric frequency-time architecture, Cyclical Calibration and Refinement (CCR) training, and Noise-Invariant Fine-Tuning (NIFT). By replacing a resource-scattering heavy frequency-domain decoder with an efficient time-domain module and using CCR to navigate multi-objective optimization, the approach achieves strong reconstruction while meeting tight budgets ($\leq 700$ MFLOPs, latency $\leq 30$ ms) and dual-rate operation (1 kbps and 6 kbps). The framework demonstrates state-of-the-art perceptual quality at 1 kbps on challenging noisy and reverberant conditions, validated through LRAC 2025 Track 1 results and ablation studies confirming the contribution of each component. The work offers a practical pathway for real-time, ultra-low-resource speech transmission in real-world networks and devices. $L_G$ is optimized as $L_G = L_{\text{mel}} + \lambda_{\text{vq}} L_{\text{vq}} + \lambda_{\text{fm}} L_{\text{fm}} + \lambda_{\text{adv}} L_{\text{adv}}$, and the CCR workflow alternates between stabilization and perceptual refinement to push performance toward the method’s theoretical limits.
Abstract
This paper presents PhoenixCodec, a comprehensive neural speech coding and decoding framework designed for extremely low-resource conditions. The proposed system integrates an optimized asymmetric frequency-time architecture, a Cyclical Calibration and Refinement (CCR) training strategy, and a noise-invariant fine-tuning procedure. Under stringent constraints - computation below 700 MFLOPs, latency less than 30 ms, and dual-rate support at 1 kbps and 6 kbps - existing methods face a trade-off between efficiency and quality. PhoenixCodec addresses these challenges by alleviating the resource scattering of conventional decoders, employing CCR to escape local optima, and enhancing robustness through noisy-sample fine-tuning. In the LRAC 2025 Challenge Track 1, the proposed system ranked third overall and demonstrated the best performance at 1 kbps in both real-world noise and reverberation and intelligibility in clean tests, confirming its effectiveness.
