Limitations of Religious Data and the Importance of the Target Domain: Towards Machine Translation for Guinea-Bissau Creole

Jacqueline Rowe; Edward Gow-Smith; Mark Hepple

Limitations of Religious Data and the Importance of the Target Domain: Towards Machine Translation for Guinea-Bissau Creole

Jacqueline Rowe, Edward Gow-Smith, Mark Hepple

TL;DR

This work addresses MT for Guinea-Bissau Creole (Kiriol), a low-resource creole with data dominated by religious texts. It introduces a ~40k-parallel-sentence dataset across Kiriol-English-Portuguese, combining religious data (Bible and JW) with a general-domain dictionary, and evaluates from-scratch Transformer models to study cross-domain transfer. Key findings show that adding a few hundred to ~600 target-domain sentences substantially improves domain-general translation, while lexical overlap with the lexifier language (Portuguese) and shared embeddings further boost performance, particularly for Kir-Por directions. The results offer practical guidance for data collection and model design in creole MT, and underscore the need for community-aware, data-rich development to enable robust language technologies for Kiriol and similar creoles.

Abstract

We introduce a new dataset for machine translation of Guinea-Bissau Creole (Kiriol), comprising around 40 thousand parallel sentences to English and Portuguese. This dataset is made up of predominantly religious data (from the Bible and texts from the Jehovah's Witnesses), but also a small amount of general domain data (from a dictionary). This mirrors the typical resource availability of many low resource languages. We train a number of transformer-based models to investigate how to improve domain transfer from religious data to a more general domain. We find that adding even 300 sentences from the target domain when training substantially improves the translation performance, highlighting the importance and need for data collection for low-resource languages, even on a small-scale. We additionally find that Portuguese-to-Kiriol translation models perform better on average than other source and target language pairs, and investigate how this relates to the morphological complexity of the languages involved and the degree of lexical overlap between creoles and lexifiers. Overall, we hope our work will stimulate research into Kiriol and into how machine translation might better support creole languages in general.

Limitations of Religious Data and the Importance of the Target Domain: Towards Machine Translation for Guinea-Bissau Creole

TL;DR

Abstract

Limitations of Religious Data and the Importance of the Target Domain: Towards Machine Translation for Guinea-Bissau Creole

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (8)