PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders

Yu Pan; Xiang Zhang; Yuguang Yang; Jixun Yao; Yanni Hu; Jianhao Ye; Hongbin Zhou; Lei Ma; Jianjun Zhao

PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders

Yu Pan, Xiang Zhang, Yuguang Yang, Jixun Yao, Yanni Hu, Jianhao Ye, Hongbin Zhou, Lei Ma, Jianjun Zhao

TL;DR

This paper proposes PSCodec, a series of neural speech codecs based on prompt encoders, comprising PSCodec-Base, PSCodec-DRL-ICT, and PSCodec-CasAN, which are capable of delivering high-performance speech reconstruction with low bandwidths.

Abstract

Neural speech codecs have recently emerged as a focal point in the fields of speech compression and generation. Despite this progress, achieving high-quality speech reconstruction under low-bitrate scenarios remains a significant challenge. In this paper, we propose PSCodec, a series of neural speech codecs based on prompt encoders, comprising PSCodec-Base, PSCodec-DRL-ICT, and PSCodec-CasAN, which are capable of delivering high-performance speech reconstruction with low bandwidths. Specifically, we first introduce PSCodec-Base, which leverages a pretrained speaker verification model-based prompt encoder (VPP-Enc) and a learnable Mel-spectrogram-based prompt encoder (MelP-Enc) to effectively disentangle and integrate voiceprint and Mel-related features in utterances. To further enhance feature utilization efficiency, we propose PSCodec-DRL-ICT, incorporating a structural similarity (SSIM) based disentangled representation loss (DRL) and an incremental continuous training (ICT) strategy. While PSCodec-DRL-ICT demonstrates impressive performance, its reliance on extensive hyperparameter tuning and multi-stage training makes it somewhat labor-intensive. To circumvent these limitations, we propose PSCodec-CasAN, utilizing an advanced cascaded attention network (CasAN) to enhance representational capacity of the entire system. Extensive experiments show that our proposed PSCodec-Base, PSCodec-DRL-ICT, and PSCodec-CasAN all significantly outperform several state-of-the-art neural codecs, exhibiting substantial improvements in both speech reconstruction quality and speaker similarity under low-bitrate conditions.

PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders

TL;DR

Abstract

Paper Structure (27 sections, 12 equations, 7 figures, 5 tables, 1 algorithm)

This paper contains 27 sections, 12 equations, 7 figures, 5 tables, 1 algorithm.

Introduction
Related Work
Discrete Neural Speech Codec.
Disentangled Representation Learning.
Methodology
PSCodec-Base
System Overview
Prompt Encoders
PSCodec-DRL-ICT
SSIM-based DRL
Incremental Continuous Training (ICT)
PSCodec-CasAN
Training Criteria
Experiments
Experimental Setups
...and 12 more sections

Figures (7)

Figure 1: The schematic of conventional neural speech codecs.
Figure 2: The architecture of our proposed PSCodec-Base training framework, comprising five fundamental components: the base encoder (Base Enc), quantizer (Quan), decoder (Dec), VPP-Enc, and MelP-Enc. Here, $X$, $X_P$, and $\tilde{X}$ represent the input speech, input prompt, and reconstructed waveform, respectively.
Figure 3: Overview of our proposed PSCodec-DRL-ICT training framework. The mutual SSIM-based DRL represents computing the SSIM scores between each pair of the three encoded features, i.e., $Z_{mel}$, $Z_{vp}$, and $Z_{q}$.
Figure 4: Detailed architecture of the proposed MelP-Enc.
Figure 5: The schematic of the proposed PSCodec-CasAN training framework. The CasAN denotes our proposed cascaded attention network.
...and 2 more figures

PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders

TL;DR

Abstract

PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders

Authors

TL;DR

Abstract

Table of Contents

Figures (7)