Table of Contents
Fetching ...

Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models

Huichan Seo, Sieun Choi, Minki Hong, Yi Zhou, Junseo Kim, Lukman Ismaila, Naome Etori, Mehul Agarwal, Zhixuan Liu, Jihie Kim, Jean Oh

TL;DR

This work tackles cultural bias in generative image models with a unified, cross-country evaluation framework that includes era-aware prompts, an 8-category/36-subcategory schema, and three I2I editing protocols across six countries. It combines standard automatic metrics, culture-aware VQA/RAG methods, and expert human judgments, and releases a complete image corpus, prompts, and configurations to enable reproducible culture-centered benchmarking. Key findings show a Global-North default under country-agnostic prompts, cultural fidelity degradation during iterative I2I edits despite stable traditional metrics, and reliance on superficial cues with partial identity preservation for Global-South targets. The study contributes an open evaluation platform and dataset, enabling robust diagnosis and tracking of cultural bias in generative image models, with implications for data curation and training objectives toward more culture-aware AI systems.

Abstract

Generative image models produce striking visuals yet often misrepresent culture. Prior work has examined cultural bias mainly in text-to-image (T2I) systems, leaving image-to-image (I2I) editors underexplored. We bridge this gap with a unified evaluation across six countries, an 8-category/36-subcategory schema, and era-aware prompts, auditing both T2I generation and I2I editing under a standardized protocol that yields comparable diagnostics. Using open models with fixed settings, we derive cross-country, cross-era, and cross-category evaluations. Our framework combines standard automatic metrics, a culture-aware retrieval-augmented VQA, and expert human judgments collected from native reviewers. To enable reproducibility, we release the complete image corpus, prompts, and configurations. Our study reveals three findings: (1) under country-agnostic prompts, models default to Global-North, modern-leaning depictions that flatten cross-country distinctions; (2) iterative I2I editing erodes cultural fidelity even when conventional metrics remain flat or improve; and (3) I2I models apply superficial cues (palette shifts, generic props) rather than era-consistent, context-aware changes, often retaining source identity for Global-South targets. These results highlight that culture-sensitive edits remain unreliable in current systems. By releasing standardized data, prompts, and human evaluation protocols, we provide a reproducible, culture-centered benchmark for diagnosing and tracking cultural bias in generative image models.

Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models

TL;DR

This work tackles cultural bias in generative image models with a unified, cross-country evaluation framework that includes era-aware prompts, an 8-category/36-subcategory schema, and three I2I editing protocols across six countries. It combines standard automatic metrics, culture-aware VQA/RAG methods, and expert human judgments, and releases a complete image corpus, prompts, and configurations to enable reproducible culture-centered benchmarking. Key findings show a Global-North default under country-agnostic prompts, cultural fidelity degradation during iterative I2I edits despite stable traditional metrics, and reliance on superficial cues with partial identity preservation for Global-South targets. The study contributes an open evaluation platform and dataset, enabling robust diagnosis and tracking of cultural bias in generative image models, with implications for data curation and training objectives toward more culture-aware AI systems.

Abstract

Generative image models produce striking visuals yet often misrepresent culture. Prior work has examined cultural bias mainly in text-to-image (T2I) systems, leaving image-to-image (I2I) editors underexplored. We bridge this gap with a unified evaluation across six countries, an 8-category/36-subcategory schema, and era-aware prompts, auditing both T2I generation and I2I editing under a standardized protocol that yields comparable diagnostics. Using open models with fixed settings, we derive cross-country, cross-era, and cross-category evaluations. Our framework combines standard automatic metrics, a culture-aware retrieval-augmented VQA, and expert human judgments collected from native reviewers. To enable reproducibility, we release the complete image corpus, prompts, and configurations. Our study reveals three findings: (1) under country-agnostic prompts, models default to Global-North, modern-leaning depictions that flatten cross-country distinctions; (2) iterative I2I editing erodes cultural fidelity even when conventional metrics remain flat or improve; and (3) I2I models apply superficial cues (palette shifts, generic props) rather than era-consistent, context-aware changes, often retaining source identity for Global-South targets. These results highlight that culture-sensitive edits remain unreliable in current systems. By releasing standardized data, prompts, and human evaluation protocols, we provide a reproducible, culture-centered benchmark for diagnosing and tracking cultural bias in generative image models.
Paper Structure (70 sections, 6 equations, 17 figures, 9 tables)

This paper contains 70 sections, 6 equations, 17 figures, 9 tables.

Figures (17)

  • Figure 1: Representative cultural biases in T2I generations across six countries. Examples include Chinese–Japanese aesthetic conflation, mis-styled Indian weddings, Kenyan wildlife stereotypes, Korean attire misidentification, Nigerian safari mislocalization, and U.S. cultural miscues in food and religious ritual. Images are from FLUX.1 [schnell] fp8 and HiDream-I1-Dev.
  • Figure 2: Overall framework overview. (a) Schema inputs: six countries, eight categories, and three era-aware prompts. (b) Experimental pipeline: T2I base generation and three I2I editing studies. (c) Multi-layered evaluation: integrating automatic, culture-aware metrics, and human evaluation.
  • Figure 3: Comparative samples for U.S. (top) and country-agnostic (bottom) prompts across two models. Left image: FLUX.1 [schnell] fp8; right image: HiDream-I1-Dev. Within each model, panels show (from left) Bride and groom, Chef, and Farmer. The close correspondence between rows illustrates the US-like default of country-agnostic prompts.
  • Figure 4: Divergence between Automated and Human Judgment in Iterative Editing. (a) CLIPScore trajectories remain largely stable or modestly increase from the base step to step 5. (b) Human Quality Score (HQS) sharply declines across all countries. This pronounced divergence highlights the failure of traditional automatic metrics to track the perceptible cultural degradation that human raters consistently penalize.
  • Figure 5: Alignment of the Culture-aware Metric with Human Judgment (All Models Averaged). (a) The agreement rate for Best Selection is high across all countries, averaging 73.8%. (b) The agreement rate for Worst Selection is consistently higher, averaging 83.7%. This high alignment demonstrates that our extended culture-aware metric successfully tracks human preference for unedited states and penalizes pronounced cultural erosion. Detailed stepwise score changes are provided in the Appendix \ref{['app:D.1']}.
  • ...and 12 more figures