Identity-Aware Large Language Models require Cultural Reasoning
Alistair Plum, Anne-Marie Lutgen, Christoph Purschke, Achim Rettinger
TL;DR
This work addresses the mismatch between large language models (LLMs) outputs and diverse cultural contexts by defining cultural reasoning (CR) and arguing that current Western-centric defaults erode trust. It proposes a four-stage, data-driven evaluation pipeline—identify domains of variation, elicit multilingual cultural descriptions, derive value statements, and post-hoc fine-tuning—to operationalize CR. The approach emphasizes validated cultural descriptions and value statements, distinguishing CR from bias mitigation, while advocating curated data, multilingual prompting, and robust benchmarks such as CulturalBench and NormAD. The goal is to enable identity-aware AI capable of context-sensitive, culturally grounded outputs, with practical impact for globally deployed systems.
Abstract
Large language models have become the latest trend in natural language processing, heavily featuring in the digital tools we use every day. However, their replies often reflect a narrow cultural viewpoint that overlooks the diversity of global users. This missing capability could be referred to as cultural reasoning, which we define here as the capacity of a model to recognise culture-specific knowledge values and social norms, and to adjust its output so that it aligns with the expectations of individual users. Because culture shapes interpretation, emotional resonance, and acceptable behaviour, cultural reasoning is essential for identity-aware AI. When this capacity is limited or absent, models can sustain stereotypes, ignore minority perspectives, erode trust, and perpetuate hate. Recent empirical studies strongly suggest that current models default to Western norms when judging moral dilemmas, interpreting idioms, or offering advice, and that fine-tuning on survey data only partly reduces this tendency. The present evaluation methods mainly report static accuracy scores and thus fail to capture adaptive reasoning in context. Although broader datasets can help, they cannot alone ensure genuine cultural competence. Therefore, we argue that cultural reasoning must be treated as a foundational capability alongside factual accuracy and linguistic coherence. By clarifying the concept and outlining initial directions for its assessment, a foundation is laid for future systems to be able to respond with greater sensitivity to the complex fabric of human culture.
