Assessing the Political Fairness of Multilingual LLMs: A Case Study based on a 21-way Multiparallel EuroParl Dataset
Paul Lerner, François Yvon
TL;DR
This work reframes political biases in multilingual LLMs as fairness in multilingual translation, proposing a speaker-affiliation aware evaluation using a new 21-EuroParl multiparallel dataset. It introduces LinkedEP-based sentence alignment across 21 languages, yielding 72,234 aligned instances and enabling cross-language, party-level analyses. By applying a Borda-count aggregation over $s$BLEU and COMET scores across 420 language pairs, the study reveals systematic translation quality advantages for major EU parties and language/translation asymmetries, highlighting potential risks in multilingual political processes. The dataset, code, and outputs are released to support reproducibility and further research in translation fairness and political NLP applications.
Abstract
The political biases of Large Language Models (LLMs) are usually assessed by simulating their answers to English surveys. In this work, we propose an alternative framing of political biases, relying on principles of fairness in multilingual translation. We systematically compare the translation quality of speeches in the European Parliament (EP), observing systematic differences with majority parties from left, center, and right being better translated than outsider parties. This study is made possible by a new, 21-way multiparallel version of EuroParl, the parliamentary proceedings of the EP, which includes the political affiliations of each speaker. The dataset consists of 1.5M sentences for a total of 40M words and 249M characters. It covers three years, 1000+ speakers, 7 countries, 12 EU parties, 25 EU committees, and hundreds of national parties.
