Table of Contents
Fetching ...

Quechua Speech Datasets in Common Voice: The Case of Puno Quechua

Elwin Huaman, Wendi Huaman, Jorge Luis Huaman, Ninfa Quispe

TL;DR

The paper addresses data scarcity for Quechua languages in speech technology and presents a case study of integrating Puno Quechua into Mozilla's Common Voice platform. It describes the onboarding and corpus-collection workflow, the distribution of Quechua data across 17 languages (191.1 hours, 86% validated) and the Puno Quechua contribution (12 hours, 77% validated). The authors propose a comprehensive research agenda tackling orthographic standardization, diverse text corpora, high-quality voice contributions, and spontaneous speech with code-switching, complemented by an ethical framework and community engagement for data sovereignty. They also emphasize practical strategies, including hybrid online-offline collection and offline tools, to broaden participation and address digital access barriers.

Abstract

Under-resourced languages, such as Quechuas, face data and resource scarcity, hindering their development in speech technology. To address this issue, Common Voice presents a crucial opportunity to foster an open and community-driven speech dataset creation. This paper examines the integration of Quechua languages into Common Voice. We detail the current 17 Quechua languages, presenting Puno Quechua (ISO 639-3: qxp) as a focused case study that includes language onboarding and corpus collection of both reading and spontaneous speech data. Our results demonstrate that Common Voice now hosts 191.1 hours of Quechua speech (86\% validated), with Puno Quechua contributing 12 hours (77\% validated), highlighting the Common Voice's potential. We further propose a research agenda addressing technical challenges, alongside ethical considerations for community engagement and indigenous data sovereignty. Our work contributes towards inclusive voice technology and digital empowerment of under-resourced language communities.

Quechua Speech Datasets in Common Voice: The Case of Puno Quechua

TL;DR

The paper addresses data scarcity for Quechua languages in speech technology and presents a case study of integrating Puno Quechua into Mozilla's Common Voice platform. It describes the onboarding and corpus-collection workflow, the distribution of Quechua data across 17 languages (191.1 hours, 86% validated) and the Puno Quechua contribution (12 hours, 77% validated). The authors propose a comprehensive research agenda tackling orthographic standardization, diverse text corpora, high-quality voice contributions, and spontaneous speech with code-switching, complemented by an ethical framework and community engagement for data sovereignty. They also emphasize practical strategies, including hybrid online-offline collection and offline tools, to broaden participation and address digital access barriers.

Abstract

Under-resourced languages, such as Quechuas, face data and resource scarcity, hindering their development in speech technology. To address this issue, Common Voice presents a crucial opportunity to foster an open and community-driven speech dataset creation. This paper examines the integration of Quechua languages into Common Voice. We detail the current 17 Quechua languages, presenting Puno Quechua (ISO 639-3: qxp) as a focused case study that includes language onboarding and corpus collection of both reading and spontaneous speech data. Our results demonstrate that Common Voice now hosts 191.1 hours of Quechua speech (86\% validated), with Puno Quechua contributing 12 hours (77\% validated), highlighting the Common Voice's potential. We further propose a research agenda addressing technical challenges, alongside ethical considerations for community engagement and indigenous data sovereignty. Our work contributes towards inclusive voice technology and digital empowerment of under-resourced language communities.
Paper Structure (23 sections, 6 figures, 1 table)

This paper contains 23 sections, 6 figures, 1 table.

Figures (6)

  • Figure 1: Language coverage by Common Voice versions.
  • Figure 2: Sentence domain distribution in Common Voice versions.
  • Figure 3: Localization efforts by Quechua communities in Common Voice platform.
  • Figure 4: Speech hours by Quechua language in Common Voice.
  • Figure 5: Comparison of a word across Southern Quechua languages.
  • ...and 1 more figures