MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis

Qian Yang; Jialong Zuo; Zhe Su; Ziyue Jiang; Mingze Li; Zhou Zhao; Feiyang Chen; Zhefeng Wang; Baoxing Huai

MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis

Qian Yang, Jialong Zuo, Zhe Su, Ziyue Jiang, Mingze Li, Zhou Zhao, Feiyang Chen, Zhefeng Wang, Baoxing Huai

TL;DR

This work addresses the need for expressive, scene-aware Mandarin TTS by releasing MSceneSpeech, a ~14.7-hour, four-scene open-source dataset with multi-speaker coverage. It introduces a robust baseline that disentangles timbre and prosody using a timbre-reference module and a prompt-based prosody mechanism, enhanced by a pretraining/fine-tuning strategy on large multi-speaker corpora. The approach employs a FastSpeech2-based architecture with Conformer-based prosody predictors and a diffusion-based decoder, enabling cross-scene style transfer and voice adaptation, as validated by MOS and ASV evaluations and comprehensive ablations. The dataset and baseline offer a practical resource for researchers and developers to build more natural, scene-specific expressive TTS systems and set a benchmark for future expressive-speech research.

Abstract

We introduce an open source high-quality Mandarin TTS dataset MSceneSpeech (Multiple Scene Speech Dataset), which is intended to provide resources for expressive speech synthesis. MSceneSpeech comprises numerous audio recordings and texts performed and recorded according to daily life scenarios. Each scenario includes multiple speakers and a diverse range of prosodic styles, making it suitable for speech synthesis that entails multi-speaker style and prosody modeling. We have established a robust baseline, through the prompting mechanism, that can effectively synthesize speech characterized by both user-specific timbre and scene-specific prosody with arbitrary text input. The open source MSceneSpeech Dataset and audio samples of our baseline are available at https://speechai-demo.github.io/MSceneSpeech/.

MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis

TL;DR

Abstract

Paper Structure (19 sections, 2 figures, 4 tables)

This paper contains 19 sections, 2 figures, 4 tables.

Introduction
Related work
Expressive Speech Synthesis Corpus
Style Transfer in Text-to-Speech
MSceneSpeech Dataset
Data recording and annotation
Data Processing
Dataset Statistics
Proposed baseline
Model overview
Prosody-Related Modules
Training and inference procedures
Experiment
Experimental setup
Model Configuration
...and 4 more sections

Figures (2)

Figure 1: Comparison of Mean of pitch, Variance of pitch, and Skew of pitch across three datasets: DidiSpeech (DS), Aishell3 (AS3), and our MSceneSpeech (MS). The statistics represent averaged metrics from individual speakers within each dataset. Each subplot focuses on a different statistical metric.
Figure 2: The overall architecture of Our Baseline. Duration, pitch, and energy are extracted from the prompt (In training: unmasked part; In inference: reference speech). It serves as conditions for their respective predictors. And losses are calculated only on the masked part.

MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis

TL;DR

Abstract

MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis

Authors

TL;DR

Abstract

Table of Contents

Figures (2)