Table of Contents
Fetching ...

PBBQ: A Persian Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models

Farhan Farsi, Shayan Bali, Fatemeh Valeh, Parsa Ghofrani, Alireza Pakniat, Kian Kashfipour, Amir H. Payberah

TL;DR

PBBQ introduces the first Persian bias QA benchmark, capturing 16 cultural topics with 233 stereotypes across 16 categories and 37,742 samples generated via human-AI collaboration. The dataset uses ambiguous and disambiguated contexts with negative and non-negative questions to measure model bias, employing accuracy, Ambiguous Bias Score, Disambiguated Bias Score, and Uncertainty metrics calculated from log-probabilities. Evaluations across eight LLMs—including open-source, closed-source, and Persian-specific models—reveal pervasive social biases in Persian LLM outputs, with disambiguation reducing but not eliminating bias and Persian-specific models sometimes aligning more closely with human biases. The work highlights the need for culturally informed data curation and context-aware evaluation to advance fair and responsible Persian language AI. PBBQ's public release will enable ongoing benchmarking and bias mitigation for Persian NLP systems across applications.

Abstract

With the increasing adoption of large language models (LLMs), ensuring their alignment with social norms has become a critical concern. While prior research has examined bias detection in various languages, there remains a significant gap in resources addressing social biases within Persian cultural contexts. In this work, we introduce PBBQ, a comprehensive benchmark dataset designed to evaluate social biases in Persian LLMs. Our benchmark, which encompasses 16 cultural categories, was developed through questionnaires completed by 250 diverse individuals across multiple demographics, in close collaboration with social science experts to ensure its validity. The resulting PBBQ dataset contains over 37,000 carefully curated questions, providing a foundation for the evaluation and mitigation of bias in Persian language models. We benchmark several open-source LLMs, a closed-source model, and Persian-specific fine-tuned models on PBBQ. Our findings reveal that current LLMs exhibit significant social biases across Persian culture. Additionally, by comparing model outputs to human responses, we observe that LLMs often replicate human bias patterns, highlighting the complex interplay between learned representations and cultural stereotypes.Upon acceptance of the paper, our PBBQ dataset will be publicly available for use in future work. Content warning: This paper contains unsafe content.

PBBQ: A Persian Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models

TL;DR

PBBQ introduces the first Persian bias QA benchmark, capturing 16 cultural topics with 233 stereotypes across 16 categories and 37,742 samples generated via human-AI collaboration. The dataset uses ambiguous and disambiguated contexts with negative and non-negative questions to measure model bias, employing accuracy, Ambiguous Bias Score, Disambiguated Bias Score, and Uncertainty metrics calculated from log-probabilities. Evaluations across eight LLMs—including open-source, closed-source, and Persian-specific models—reveal pervasive social biases in Persian LLM outputs, with disambiguation reducing but not eliminating bias and Persian-specific models sometimes aligning more closely with human biases. The work highlights the need for culturally informed data curation and context-aware evaluation to advance fair and responsible Persian language AI. PBBQ's public release will enable ongoing benchmarking and bias mitigation for Persian NLP systems across applications.

Abstract

With the increasing adoption of large language models (LLMs), ensuring their alignment with social norms has become a critical concern. While prior research has examined bias detection in various languages, there remains a significant gap in resources addressing social biases within Persian cultural contexts. In this work, we introduce PBBQ, a comprehensive benchmark dataset designed to evaluate social biases in Persian LLMs. Our benchmark, which encompasses 16 cultural categories, was developed through questionnaires completed by 250 diverse individuals across multiple demographics, in close collaboration with social science experts to ensure its validity. The resulting PBBQ dataset contains over 37,000 carefully curated questions, providing a foundation for the evaluation and mitigation of bias in Persian language models. We benchmark several open-source LLMs, a closed-source model, and Persian-specific fine-tuned models on PBBQ. Our findings reveal that current LLMs exhibit significant social biases across Persian culture. Additionally, by comparing model outputs to human responses, we observe that LLMs often replicate human bias patterns, highlighting the complex interplay between learned representations and cultural stereotypes.Upon acceptance of the paper, our PBBQ dataset will be publicly available for use in future work. Content warning: This paper contains unsafe content.
Paper Structure (31 sections, 9 equations, 10 figures, 10 tables)

This paper contains 31 sections, 9 equations, 10 figures, 10 tables.

Figures (10)

  • Figure 1: Overview of dataset construction process, which involves 4 stages: selecting bias topics, extracting stereotypes, generating contexts from templates, and creating corresponding negative/non-negative of questions.
  • Figure 2: An example from the PBBQ dataset. The green box highlights the bias topic and extracted stereotype for this instance. The blue box presents the context templates along with the placeholders used to populate them. The red box illustrates the corresponding negative and non-negative questions derived from the contexts. The purple box displays the answer for each type of question based on the provided scenario.
  • Figure 3: Uncertainty score box plot (0-1) across models on the PBBQ dataset for both ambiguous and disambiguated contexts.
  • Figure 4: Gender distribution of participants: 140 male, 110 female
  • Figure 5: Age distribution of participants: 85 were aged 18–24, 68 were 25–34 , 52 were 35–44, 27 were 45–54, and 18 were 55+
  • ...and 5 more figures