Table of Contents
Fetching ...

ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models

Emily Chang, Niyati Bafna

TL;DR

ChiKhaPo tackles language inequity by delivering a massively multilingual lexical benchmark spanning 2700+ languages. It defines 8 subtasks across two evaluation directions to probe word-level lexical comprehension and generation, using lexicons, monolingual data, and bitext to maximize language coverage. The study evaluates six multilingual LLMs, revealing that comprehension directions generally outperform generation, with Indo-European languages typically scoring higher than underrepresented families and WT showing strong correlation with MT benchmarks. Overall, ChiKhaPo demonstrates both the feasibility and the need for broad multilingual lexical evaluation to drive improvements in low-resource languages and guides future resource collection and benchmarking efforts.

Abstract

Existing benchmarks for large language models (LLMs) are largely restricted to high- or mid-resource languages, and often evaluate performance on higher-order tasks in reasoning and generation. However, plenty of evidence points to the fact that LLMs lack basic linguistic competence in the vast majority of the world's 3800+ written languages. We introduce ChiKhaPo, consisting of 8 subtasks of varying difficulty designed to evaluate the lexical comprehension and generation abilities of generative models. ChiKhaPo draws on existing lexicons, monolingual data, and bitext, and provides coverage for 2700+ languages for 2 subtasks, surpassing any existing benchmark in terms of language coverage. We further show that 6 SOTA models struggle on our benchmark, and discuss the factors contributing to performance scores, including language family, language resourcedness, task, and comprehension versus generation directions. With ChiKhaPo, we hope to enable and encourage the massively multilingual benchmarking of LLMs.

ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models

TL;DR

ChiKhaPo tackles language inequity by delivering a massively multilingual lexical benchmark spanning 2700+ languages. It defines 8 subtasks across two evaluation directions to probe word-level lexical comprehension and generation, using lexicons, monolingual data, and bitext to maximize language coverage. The study evaluates six multilingual LLMs, revealing that comprehension directions generally outperform generation, with Indo-European languages typically scoring higher than underrepresented families and WT showing strong correlation with MT benchmarks. Overall, ChiKhaPo demonstrates both the feasibility and the need for broad multilingual lexical evaluation to drive improvements in low-resource languages and guides future resource collection and benchmarking efforts.

Abstract

Existing benchmarks for large language models (LLMs) are largely restricted to high- or mid-resource languages, and often evaluate performance on higher-order tasks in reasoning and generation. However, plenty of evidence points to the fact that LLMs lack basic linguistic competence in the vast majority of the world's 3800+ written languages. We introduce ChiKhaPo, consisting of 8 subtasks of varying difficulty designed to evaluate the lexical comprehension and generation abilities of generative models. ChiKhaPo draws on existing lexicons, monolingual data, and bitext, and provides coverage for 2700+ languages for 2 subtasks, surpassing any existing benchmark in terms of language coverage. We further show that 6 SOTA models struggle on our benchmark, and discuss the factors contributing to performance scores, including language family, language resourcedness, task, and comprehension versus generation directions. With ChiKhaPo, we hope to enable and encourage the massively multilingual benchmarking of LLMs.
Paper Structure (82 sections, 12 equations, 15 figures, 24 tables)

This paper contains 82 sections, 12 equations, 15 figures, 24 tables.

Figures (15)

  • Figure 1: ChiKhaPo evaluates basic lexical competence with several tasks, covering an order of magnitude more languages than existing multilingual benchmarks.
  • Figure 2: Model scores across subtasks, with std. deviation over languages. Best performing model is highlighted.
  • Figure 3: We compute the score of a language family as the average of its constituent languages, with the best-performing language family highlighted. Error bars represent the standard deviation within the language family. The Indo-European family has consistently higher scores than other families.
  • Figure 4: Comparison of the number of Wikipedia documents— a proxy for resource level— and language performance for the task WT. See \ref{['subsec:resourceness']} for other tasks.
  • Figure 5: WT scores are strongly correlated with sentence-level MT BLEU scores.
  • ...and 10 more figures