Table of Contents
Fetching ...

I Am Aligned, But With Whom? MENA Values Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMs

Pardis Sadat Zahraei, Ehsaneddin Asgari

TL;DR

The paper tackles the misalignment of large language models with the diverse cultural values of the MENA region and the multilingual biases that emerge across languages. It introduces MENAValues, a benchmark built from World Values Survey Wave 7 and the Arab Opinion Index, comprising 864 questions across 16 countries to assess cultural alignment under multiple framing and language conditions. The authors reveal three robust phenomena—Cross-Lingual Value Shifts, Reasoning-Induced Degradation, and Logit Leakage—and show that models often collapse into language-based clusters, undermining genuine cross-cultural understanding. The work provides a scalable evaluation framework with empirical methods (NVAS, FCS, CLCS, SPD, token-probability analysis, PCA) to diagnose alignment failures and guide the development of more culturally inclusive AI, with implications for multilingual deployment in diverse regions and beyond.

Abstract

We introduce MENAValues, a novel benchmark designed to evaluate the cultural alignment and multilingual biases of large language models (LLMs) with respect to the beliefs and values of the Middle East and North Africa (MENA) region, an underrepresented area in current AI evaluation efforts. Drawing from large-scale, authoritative human surveys, we curate a structured dataset that captures the sociocultural landscape of MENA with population-level response distributions from 16 countries. To probe LLM behavior, we evaluate diverse models across multiple conditions formed by crossing three perspective framings (neutral, personalized, and third-person/cultural observer) with two language modes (English and localized native languages: Arabic, Persian, Turkish). Our analysis reveals three critical phenomena: "Cross-Lingual Value Shifts" where identical questions yield drastically different responses based on language, "Reasoning-Induced Degradation" where prompting models to explain their reasoning worsens cultural alignment, and "Logit Leakage" where models refuse sensitive questions while internal probabilities reveal strong hidden preferences. We further demonstrate that models collapse into simplistic linguistic categories when operating in native languages, treating diverse nations as monolithic entities. MENAValues offers a scalable framework for diagnosing cultural misalignment, providing both empirical insights and methodological tools for developing more culturally inclusive AI.

I Am Aligned, But With Whom? MENA Values Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMs

TL;DR

The paper tackles the misalignment of large language models with the diverse cultural values of the MENA region and the multilingual biases that emerge across languages. It introduces MENAValues, a benchmark built from World Values Survey Wave 7 and the Arab Opinion Index, comprising 864 questions across 16 countries to assess cultural alignment under multiple framing and language conditions. The authors reveal three robust phenomena—Cross-Lingual Value Shifts, Reasoning-Induced Degradation, and Logit Leakage—and show that models often collapse into language-based clusters, undermining genuine cross-cultural understanding. The work provides a scalable evaluation framework with empirical methods (NVAS, FCS, CLCS, SPD, token-probability analysis, PCA) to diagnose alignment failures and guide the development of more culturally inclusive AI, with implications for multilingual deployment in diverse regions and beyond.

Abstract

We introduce MENAValues, a novel benchmark designed to evaluate the cultural alignment and multilingual biases of large language models (LLMs) with respect to the beliefs and values of the Middle East and North Africa (MENA) region, an underrepresented area in current AI evaluation efforts. Drawing from large-scale, authoritative human surveys, we curate a structured dataset that captures the sociocultural landscape of MENA with population-level response distributions from 16 countries. To probe LLM behavior, we evaluate diverse models across multiple conditions formed by crossing three perspective framings (neutral, personalized, and third-person/cultural observer) with two language modes (English and localized native languages: Arabic, Persian, Turkish). Our analysis reveals three critical phenomena: "Cross-Lingual Value Shifts" where identical questions yield drastically different responses based on language, "Reasoning-Induced Degradation" where prompting models to explain their reasoning worsens cultural alignment, and "Logit Leakage" where models refuse sensitive questions while internal probabilities reveal strong hidden preferences. We further demonstrate that models collapse into simplistic linguistic categories when operating in native languages, treating diverse nations as monolithic entities. MENAValues offers a scalable framework for diagnosing cultural misalignment, providing both empirical insights and methodological tools for developing more culturally inclusive AI.
Paper Structure (78 sections, 4 equations, 42 figures, 3 tables)

This paper contains 78 sections, 4 equations, 42 figures, 3 tables.

Figures (42)

  • Figure 1: Systematic Value Inconsistency in LLMs: A Multi-Dimensional Analysis of Alignment Failures. This figure reveals how LLMs exhibit inconsistencies when responding to value-based questions, demonstrating three critical dimensions of misalignment. Cross-Lingual Value Shift shows how identical questions yield contradictory responses across languages (Arabic vs. English), suggesting cultural bias encoding rather than consistent moral reasoning. Prompt-Sensitive Misalignment reveals how different framings (Neutral, Persona, Observer) elicit conflicting stances, indicating unstable value representation. Logit Leakage exposes how surface-level safety responses can mask concerning internal preferences, with high-confidence hidden biases contradicting stated positions.
  • Figure 2: Core Dimensions of the MENAValues Dataset. The dataset is structured around four major pillars: (1) Social & Cultural Identity, (2) Economic Dimensions, (3) Governance & Political Systems, and (4) Individual Wellbeing & Development. Each category is illustrated with survey questions and average responses from representative MENA countries. Note: The countries shown are illustrative examples; all benchmark questions are posed to LLMs for all 16 countries in our dataset, regardless of original human data availability.
  • Figure 3: Principal Component Analysis of WVS-7 countries.
  • Figure 4: Distributional similarity heatmaps showing Jensen-Shannon Divergence-based similarity scores (0-1 scale, where 1 indicates identical distributions) between country pairs across thematic categories.
  • Figure 5: Principal Component Analysis of AOI countries.
  • ...and 37 more figures