ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation

Ruibo Fu; Xin Qi; Zhengqi Wen; Jianhua Tao; Tao Wang; Chunyu Qiang; Zhiyong Wang; Yi Lu; Xiaopeng Wang; Shuchen Shi; Yukun Liu; Xuefei Liu; Shuai Zhang

ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation

Ruibo Fu, Xin Qi, Zhengqi Wen, Jianhua Tao, Tao Wang, Chunyu Qiang, Zhiyong Wang, Yi Lu, Xiaopeng Wang, Shuchen Shi, Yukun Liu, Xuefei Liu, Shuai Zhang

TL;DR

This paper tackles the problem of speaker adaptation in Text-to-Speech with limited target references, where traditional fine-tuning and speaker-encoder approaches struggle with speaker representation and overfitting. It introduces Agile Speaker Representation Reinforcement Learning (ASRRL), a method that refines speaker embeddings via reinforcement learning without altering the core TTS model, using two scenario-specific action strategies (SS and FS) and a multi-dimensional fusion reward to balance speaker similarity, speech quality, and intelligibility. Key contributions include a prior-guided state representation, a two-scenario action design, and a fusion-based reward that stabilizes optimization; extensive experiments on LibriTTS and VCTK across VITS and Grad-TTS backbones show superior speaker similarity and robust quality, especially when reference speech is scarce. The results demonstrate ASRRL’s extensibility and potential for real-world applications in personalized TTS, with implications for applying reinforcement learning to broader audio generation tasks.

Abstract

Speaker adaptation, which involves cloning voices from unseen speakers in the Text-to-Speech task, has garnered significant interest due to its numerous applications in multi-media fields. Despite recent advancements, existing methods often struggle with inadequate speaker representation accuracy and overfitting, particularly in limited reference speeches scenarios. To address these challenges, we propose an Agile Speaker Representation Reinforcement Learning strategy to enhance speaker similarity in speaker adaptation tasks. ASRRL is the first work to apply reinforcement learning to improve the modeling accuracy of speaker embeddings in speaker adaptation, addressing the challenge of decoupling voice content and timbre. Our approach introduces two action strategies tailored to different reference speeches scenarios. In the single-sentence scenario, a knowledge-oriented optimal routine searching RL method is employed to expedite the exploration and retrieval of refinement information on the fringe of speaker representations. In the few-sentence scenario, we utilize a dynamic RL method to adaptively fuse reference speeches, enhancing the robustness and accuracy of speaker modeling. To achieve optimal results in the target domain, a multi-scale fusion scoring mechanism based reward model that evaluates speaker similarity, speech quality, and intelligibility across three dimensions is proposed, ensuring that improvements in speaker similarity do not compromise speech quality or intelligibility. The experimental results on the LibriTTS and VCTK datasets within mainstream TTS frameworks demonstrate the extensibility and generalization capabilities of the proposed ASRRL method. The results indicate that the ASRRL method significantly outperforms traditional fine-tuning approaches, achieving higher speaker similarity and better overall speech quality with limited reference speeches.

ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation

TL;DR

Abstract

Paper Structure (30 sections, 2 equations, 5 figures, 9 tables)

This paper contains 30 sections, 2 equations, 5 figures, 9 tables.

Introduction
Related Work
Speaker Adaptation
Reinforcement learning
Method
The framework of ASSRL
Deep generative models
Agile Speaker Representation Reinforcement Learning
Prior-Guided State Representation
Two-Scenario Action Design
Multi-Dimensional Fusion Reward Design
Experiment and Analysis
Experimental Setup
Dataset
Task
...and 15 more sections

Figures (5)

Figure 1: The framework of ASSRL: (1)Extract the speaker representation embedding $e$ from the reference speech $s_r$. Additionally, extract text features $f_t$ using a pre-trained Bert model. These text features and speaker embeddings form the state representation for RL. A pre-trained TTS model $F$ synthesizes the text features and speaker embeddings into speech $s_s$. The reward model then evaluates the current generation results using a fusion scoring mechanism. Based on the scenario, the RL agent selects different actions to alter the speaker embedding information $e$ and obtain the next state. (2)When there are a few reference speeches, the initial state's speaker embedding part is represented by the mean of the speaker embeddings extracted by the encoder. (3)During inference, the trained policy is utilized to modify the speaker embedding information $e$, resulting in an enhanced speaker similarity denoted by $e^*$.
Figure 2: The process of Agile Speaker Representation Reinforcement Learning
Figure 3: State Representation. Text features $f_t$. Speaker representation embedding $e$. Separator $<sep>$.
Figure 4: Two-Scenario Action Design. In the SS scenario, the output action is a refinement with the same dimensions as $e$. In the FS scenario, the output action is the fusion ratio for a few reference speeches.
Figure 5: The results of hyperparameter changes in reinforcement learning. (a) and (b) show the impact of changes in gamma values on similarity during reinforcement learning. (c) and (d) show the impact of changes in action scaling factor. (e) and (f) show the impact of changes in the number of exploration steps.

ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation

TL;DR

Abstract

ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation

Authors

TL;DR

Abstract

Table of Contents

Figures (5)