Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data

Hitoshi Suda; Aya Watanabe; Shinnosuke Takamichi

Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data

Hitoshi Suda, Aya Watanabe, Shinnosuke Takamichi

TL;DR

This study addresses who finds a voice attractive by constructing CocoNut-Humoresque, a large-scale in-the-wild corpus of 1800 speech segments rated for likability by 885 listeners, with rich listener and speaker attributes. It analyzes gender- and age-related biases in likability, and investigates how acoustic features such as the fundamental frequency $F_0$ and speaker embeddings ($x$-vectors) relate to perceived attractiveness, using MOS analyses and $t$-SNE visualizations. Key findings include systematic gender and age biases in likability, and evidence that $F_0$ and $x$-vector embeddings partly predict likability while still leaving substantial influence from other factors. The dataset and findings provide a valuable resource for designing voice systems and understanding listener diversity in voice preference.

Abstract

This paper introduces CocoNut-Humoresque, an open-source large-scale speech likability corpus that includes speech segments and their per-listener likability scores. Evaluating voice likability is essential to designing preferable voices for speech systems, such as dialogue or announcement systems. In this study, we let 885 listeners rate 1800 speech segments of a wide range of speakers regarding their likability. When constructing the corpus, we also collected the multiple speaker attributes: genders, ages, and favorite YouTube videos. Therefore, the corpus enables the large-scale statistical analysis of voice likability regarding both speaker and listener factors. This paper describes the construction methodology and preliminary data analysis to reveal the gender and age biases in voice likability. In addition, the relationship between the likability and two acoustic features, the fundamental frequencies and the x-vectors of given utterances, is also investigated.

Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data

TL;DR

and speaker embeddings (

-vectors) relate to perceived attractiveness, using MOS analyses and

-SNE visualizations. Key findings include systematic gender and age biases in likability, and evidence that

and

-vector embeddings partly predict likability while still leaving substantial influence from other factors. The dataset and findings provide a valuable resource for designing voice systems and understanding listener diversity in voice preference.

Abstract

Paper Structure (13 sections, 7 figures, 2 tables, 1 algorithm)

This paper contains 13 sections, 7 figures, 2 tables, 1 algorithm.

Introduction
CocoNut-Humoresque: A large-scale speech likability corpus
Voice materials
Corpus design
Listener attributes
Analysis 1: Gender and age biases
Gender biases
Age biases
Analysis 2: Sample-by-sample analysis
Likable voices only for males or females
Relationship between likability and x-vectors
Conclusions
Acknowledgements

Figures (7)

Figure 1: Visualization of the contents of the 80 evaluation subsets in the corpus. The columns show the speech segments, and the rows show the subsets.
Figure 2: Violin plots of the MOSs by genders of listeners and speakers. The left of each column's name denotes the listener's gender, and the right part denotes the speaker's gender. M, F, and A denote males, females, and whole data, respectively.
Figure 3: Violin plots of the MOSs by the listener's age. Asterisks denote significant differences with $p<0.01$.
Figure 4: Violin plots of the MOSs by the listener's gender and age. Asterisks denote significant differences with $p<0.01$.
Figure 5: Relationship between scores by male and female listeners. The samples in the upper left part are likable, especially for female listeners, and vice versa. The points (a) and (b) represent the samples with the most divided opinions between male and female listeners. The $F_0$ means were calculated using Crepe's full modelKim2018-ij.
...and 2 more figures

Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data

TL;DR

Abstract

Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data

Authors

TL;DR

Abstract

Table of Contents

Figures (7)