Table of Contents
Fetching ...

Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps

Ahmed Alzubaidi, Shaikha Alsuwaidi, Basma El Amel Boussaha, Leen AlQadi, Omar Alkaabi, Mohammed Alyafeai, Hamza Alobeidli, Hakim Hacid

TL;DR

This survey systematically analyzes 40+ Arabic LLM benchmarks, introducing a four-category taxonomy (Knowledge, NLP Tasks, Culture and Dialects, Target-Specific) to organize evaluation datasets. It compares native-, translation-, and synthetic-data approaches, highlighting cultural alignment challenges and the emergence of LLM-as-Judge-driven benchmarks. The authors identify progress toward unified, dialect-aware evaluation (via benchmarks like LAraBench, BALSAM, and ORCA) alongside critical gaps in temporal reasoning, multi-turn dialogue, and reproducibility. They provide concrete recommendations and a community resource repository to standardize evaluation practices and foster robust, culturally authentic Arabic LLM assessment.

Abstract

This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and specialized capabilities. We propose a taxonomy organizing benchmarks into four categories: Knowledge, NLP Tasks, Culture and Dialects, and Target-Specific evaluations. Our analysis reveals significant progress in benchmark diversity while identifying critical gaps: limited temporal evaluation, insufficient multi-turn dialogue assessment, and cultural misalignment in translated datasets. We examine three primary approaches: native collection, translation, and synthetic generation discussing their trade-offs regarding authenticity, scale, and cost. This work serves as a comprehensive reference for Arabic NLP researchers, providing insights into benchmark methodologies, reproducibility standards, and evaluation metrics while offering recommendations for future development.

Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods, and Gaps

TL;DR

This survey systematically analyzes 40+ Arabic LLM benchmarks, introducing a four-category taxonomy (Knowledge, NLP Tasks, Culture and Dialects, Target-Specific) to organize evaluation datasets. It compares native-, translation-, and synthetic-data approaches, highlighting cultural alignment challenges and the emergence of LLM-as-Judge-driven benchmarks. The authors identify progress toward unified, dialect-aware evaluation (via benchmarks like LAraBench, BALSAM, and ORCA) alongside critical gaps in temporal reasoning, multi-turn dialogue, and reproducibility. They provide concrete recommendations and a community resource repository to standardize evaluation practices and foster robust, culturally authentic Arabic LLM assessment.

Abstract

This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and specialized capabilities. We propose a taxonomy organizing benchmarks into four categories: Knowledge, NLP Tasks, Culture and Dialects, and Target-Specific evaluations. Our analysis reveals significant progress in benchmark diversity while identifying critical gaps: limited temporal evaluation, insufficient multi-turn dialogue assessment, and cultural misalignment in translated datasets. We examine three primary approaches: native collection, translation, and synthetic generation discussing their trade-offs regarding authenticity, scale, and cost. This work serves as a comprehensive reference for Arabic NLP researchers, providing insights into benchmark methodologies, reproducibility standards, and evaluation metrics while offering recommendations for future development.
Paper Structure (24 sections, 3 figures, 2 tables)

This paper contains 24 sections, 3 figures, 2 tables.

Figures (3)

  • Figure 1: Taxonomy of Arabic Benchmarks.
  • Figure 2: Arabic Datasets Sizes Across Categories. Size is log-scaled.
  • Figure 3: Timeline of Arabic benchmark releases (2019-2025) showing dramatic acceleration, with 82% of the benchmarks released in 2024-2025.