Cost Analysis of Human-corrected Transcription for Predominately Oral Languages
Yacouba Diarra, Nouhoum Souleymane Coulibaly, Michael Leventhal
TL;DR
This paper tackles the cost of creating NLP resources for Predominately Oral Languages by conducting a one-month field study in Bambara with ten native-level transcribers who post-edit ASR-generated transcripts for 53 hours of speech. The authors implement a first-review transcription workflow using Label Studio, VAD-based segmentation, and pre-labeled ASR outputs to measure human labor, achieving a labor cost of about 30 hours per hour of annotated audio under ideal conditions and roughly 36 hours per hour in field conditions. The study provides a concrete, empirically grounded baseline for budgeting and planning data collection for low-literacy POLs, highlighting the substantial cognitive and logistic demands that exceed those typical for higher-literacy languages. The findings have practical implications for resource allocation, pipeline design, and the evaluation of synthetic data approaches for building robust speech technologies in POLs. By clarifying the labor requirements, the work informs policy and practice for scalable NLP resource development in low-resource, predominantly oral language communities.
Abstract
Creating speech datasets for low-resource languages is a critical yet poorly understood challenge, particularly regarding the actual cost in human labor. This paper investigates the time and complexity required to produce high-quality annotated speech data for a subset of low-resource languages, low literacy Predominately Oral Languages, focusing on Bambara, a Manding language of Mali. Through a one-month field study involving ten transcribers with native proficiency, we analyze the correction of ASR-generated transcriptions of 53 hours of Bambara voice data. We report that it takes, on average, 30 hours of human labor to accurately transcribe one hour of speech data under laboratory conditions and 36 hours under field conditions. The study provides a baseline and practical insights for a large class of languages with comparable profiles undertaking the creation of NLP resources.
