Table of Contents
Fetching ...

Are the LLMs Capable of Maintaining at Least the Language Genus?

Sandra Mitrović, David Kletz, Ljiljana Dolamic, Fabio Rinaldi

TL;DR

This work investigates whether LLMs exhibit genus-level sensitivity in multilingual behavior using a genus-centric extension of the MultiQ benchmark. By mapping languages to 47 genera across 21 families and evaluating both fidelity and cross-genus knowledge transfer, the study reveals that genus-level effects exist but are strongly shaped by training data availability; different model families employ distinct multilingual strategies, and resource imbalance in training data largely governs multilingual performance. The authors introduce FidelityScore and SwitchScore metrics to quantify within-genus consistency and cross-genus knowledge transfer, respectively, demonstrating higher intra-genus stability (80–90%) compared to cross-genus transfers (40–50%) and notable asymmetries across language directions. Overall, the work highlights that genealogical structure is encoded to some extent in LLMs, but data distribution remains the dominant factor driving multilingual behavior, with important implications for dataset design, model training, and evaluation of multilingual capabilities. These findings suggest a need to balance training data across genera to improve robust multilingual performance across diverse languages.

Abstract

Large Language Models (LLMs) display notable variation in multilingual behavior, yet the role of genealogical language structure in shaping this variation remains underexplored. In this paper, we investigate whether LLMs exhibit sensitivity to linguistic genera by extending prior analyses on the MultiQ dataset. We first check if models prefer to switch to genealogically related languages when prompt language fidelity is not maintained. Next, we investigate whether knowledge consistency is better preserved within than across genera. We show that genus-level effects are present but strongly conditioned by training resource availability. We further observe distinct multilingual strategies across LLMs families. Our findings suggest that LLMs encode aspects of genus-level structure, but training data imbalances remain the primary factor shaping their multilingual performance.

Are the LLMs Capable of Maintaining at Least the Language Genus?

TL;DR

This work investigates whether LLMs exhibit genus-level sensitivity in multilingual behavior using a genus-centric extension of the MultiQ benchmark. By mapping languages to 47 genera across 21 families and evaluating both fidelity and cross-genus knowledge transfer, the study reveals that genus-level effects exist but are strongly shaped by training data availability; different model families employ distinct multilingual strategies, and resource imbalance in training data largely governs multilingual performance. The authors introduce FidelityScore and SwitchScore metrics to quantify within-genus consistency and cross-genus knowledge transfer, respectively, demonstrating higher intra-genus stability (80–90%) compared to cross-genus transfers (40–50%) and notable asymmetries across language directions. Overall, the work highlights that genealogical structure is encoded to some extent in LLMs, but data distribution remains the dominant factor driving multilingual behavior, with important implications for dataset design, model training, and evaluation of multilingual capabilities. These findings suggest a need to balance training data across genera to improve robust multilingual performance across diverse languages.

Abstract

Large Language Models (LLMs) display notable variation in multilingual behavior, yet the role of genealogical language structure in shaping this variation remains underexplored. In this paper, we investigate whether LLMs exhibit sensitivity to linguistic genera by extending prior analyses on the MultiQ dataset. We first check if models prefer to switch to genealogically related languages when prompt language fidelity is not maintained. Next, we investigate whether knowledge consistency is better preserved within than across genera. We show that genus-level effects are present but strongly conditioned by training resource availability. We further observe distinct multilingual strategies across LLMs families. Our findings suggest that LLMs encode aspects of genus-level structure, but training data imbalances remain the primary factor shaping their multilingual performance.
Paper Structure (27 sections, 3 equations, 12 figures, 3 tables)

This paper contains 27 sections, 3 equations, 12 figures, 3 tables.

Figures (12)

  • Figure 1: Fidelity scores across genera (existing in MultiQ) and models (existing in MultiQ + Apertus).
  • Figure 2: Genus-level fidelity across models. For each representative genus, we report the proportion of model outputs that remain within the same genus as the prompt language.
  • Figure 3: Genus-level output distribution by model. For each prompt genus, we indicate the genus of the model’s generated response. Remaining models are reported in Appendix \ref{['app:detail_output_model']}.
  • Figure 4: Switchscores across models and thresholds. Each subfigure shows the switchscore distribution for one model at the specified threshold. As can be seen from red/blue-column patterns, the performance critically depends on the target genus.
  • Figure 5: Genus-level output distribution by model. For each prompt genus, we indicate the genus of the model’s generated response.
  • ...and 7 more figures