Table of Contents
Fetching ...

StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks

Jingyue Huang, Qihui Yang, Fei Yueh Chen, Julian McAuley, Randal Leistikow, Perry R. Cook, Yongyi Zang

TL;DR

StylePitcher addresses the need for expressive, singer-specific pitch curves that generalize across singing tasks. It introduces a rectified flow matching framework for conditional infilling, learning style from reference audio while staying aligned to the target melody. The method achieves superior style similarity and audio quality with pitch accuracy comparable to task-specific baselines, without retraining for each application, enabling plug-and-play deployment for automatic pitch correction, zero-shot SVS with style transfer, and style-informed SVC. This approach has practical impact by broadening the applicability of expressive pitch modeling across diverse singing systems.

Abstract

Existing pitch curve generators face two main challenges: they often neglect singer-specific expressiveness, reducing their ability to capture individual singing styles. And they are typically developed as auxiliary modules for specific tasks such as pitch correction, singing voice synthesis, or voice conversion, which restricts their generalization capability. We propose StylePitcher, a general-purpose pitch curve generator that learns singer style from reference audio while preserving alignment with the intended melody. Built upon a rectified flow matching architecture, StylePitcher flexibly incorporates symbolic music scores and pitch context as conditions for generation, and can seamlessly adapt to diverse singing tasks without retraining. Objective and subjective evaluations across various singing tasks demonstrate that StylePitcher improves style similarity and audio quality while maintaining pitch accuracy comparable to task-specific baselines.

StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks

TL;DR

StylePitcher addresses the need for expressive, singer-specific pitch curves that generalize across singing tasks. It introduces a rectified flow matching framework for conditional infilling, learning style from reference audio while staying aligned to the target melody. The method achieves superior style similarity and audio quality with pitch accuracy comparable to task-specific baselines, without retraining for each application, enabling plug-and-play deployment for automatic pitch correction, zero-shot SVS with style transfer, and style-informed SVC. This approach has practical impact by broadening the applicability of expressive pitch modeling across diverse singing systems.

Abstract

Existing pitch curve generators face two main challenges: they often neglect singer-specific expressiveness, reducing their ability to capture individual singing styles. And they are typically developed as auxiliary modules for specific tasks such as pitch correction, singing voice synthesis, or voice conversion, which restricts their generalization capability. We propose StylePitcher, a general-purpose pitch curve generator that learns singer style from reference audio while preserving alignment with the intended melody. Built upon a rectified flow matching architecture, StylePitcher flexibly incorporates symbolic music scores and pitch context as conditions for generation, and can seamlessly adapt to diverse singing tasks without retraining. Objective and subjective evaluations across various singing tasks demonstrate that StylePitcher improves style similarity and audio quality while maintaining pitch accuracy comparable to task-specific baselines.
Paper Structure (14 sections, 5 equations, 2 figures, 2 tables)

This paper contains 14 sections, 5 equations, 2 figures, 2 tables.

Figures (2)

  • Figure 1: Illustration of the methods. The unvoiced condition is omitted for clarity. Subscripts off and in denote features from off-key and in-key singing in the APC task; ref and tgt refer to reference and target content for SVS and SVC tasks.
  • Figure 2: Samples for three singing tasks. StylePitcher (red) captures singing styles better from input curves (blue) than baselines (green), such as pitch slides (a) and vibrato (b&c).