Table of Contents
Fetching ...

Race and Gender in LLM-Generated Personas: A Large-Scale Audit of 41 Occupations

Ilona van der Linden, Sahana Kumar, Arnav Dixit, Aadi Sudan, Smruthi Danda, David C. Anastasiu, Kai Lukoff

TL;DR

This study audits over 1.5 million LLM-generated occupational personas across 41 U.S. occupations to assess race and gender representation against BLS baselines. Using four diverse models and a regression framework with $\alpha$ (systematic shift) and $\beta$ (stereotype amplification), the authors find consistent distortions: White and Black workers are underrepresented while Hispanic and Asian workers are overrepresented, with extreme examples such as nearly all Housekeepers being Hispanic. The analysis reveals cross-model convergence on biased patterns but model-specific offsets, underscoring that provider choice materially shapes visibility and representation. The work advocates for design interventions in persona tools—including explicit baselines, multi-pesona sets, and process transparency—to govern representation and foster accountability in AI systems.

Abstract

Generative AI tools are increasingly used to create portrayals of people in occupations, raising concerns about how race and gender are represented. We conducted a large-scale audit of over 1.5 million occupational personas across 41 U.S. occupations, generated by four large language models with different AI safety commitments and countries of origin (U.S., China, France). Compared with Bureau of Labor Statistics data, we find two recurring patterns: systematic shifts, where some groups are consistently under- or overrepresented, and stereotype exaggeration, where existing demographic skews are amplified. On average, White (--31pp) and Black (--9pp) workers are underrepresented, while Hispanic (+17pp) and Asian (+12pp) workers are overrepresented. These distortions can be extreme: for example, across all four models, Housekeepers are portrayed as nearly 100\% Hispanic, while Black workers are erased from many occupations. For HCI, these findings show provider choice materially changes who is visible, motivating model-specific audits and accountable design practices.

Race and Gender in LLM-Generated Personas: A Large-Scale Audit of 41 Occupations

TL;DR

This study audits over 1.5 million LLM-generated occupational personas across 41 U.S. occupations to assess race and gender representation against BLS baselines. Using four diverse models and a regression framework with (systematic shift) and (stereotype amplification), the authors find consistent distortions: White and Black workers are underrepresented while Hispanic and Asian workers are overrepresented, with extreme examples such as nearly all Housekeepers being Hispanic. The analysis reveals cross-model convergence on biased patterns but model-specific offsets, underscoring that provider choice materially shapes visibility and representation. The work advocates for design interventions in persona tools—including explicit baselines, multi-pesona sets, and process transparency—to govern representation and foster accountability in AI systems.

Abstract

Generative AI tools are increasingly used to create portrayals of people in occupations, raising concerns about how race and gender are represented. We conducted a large-scale audit of over 1.5 million occupational personas across 41 U.S. occupations, generated by four large language models with different AI safety commitments and countries of origin (U.S., China, France). Compared with Bureau of Labor Statistics data, we find two recurring patterns: systematic shifts, where some groups are consistently under- or overrepresented, and stereotype exaggeration, where existing demographic skews are amplified. On average, White (--31pp) and Black (--9pp) workers are underrepresented, while Hispanic (+17pp) and Asian (+12pp) workers are overrepresented. These distortions can be extreme: for example, across all four models, Housekeepers are portrayed as nearly 100\% Hispanic, while Black workers are erased from many occupations. For HCI, these findings show provider choice materially changes who is visible, motivating model-specific audits and accountable design practices.
Paper Structure (42 sections, 7 figures, 5 tables)

This paper contains 42 sections, 7 figures, 5 tables.

Figures (7)

  • Figure 1: Conceptual illustration of how different bias patterns can appear when comparing large language model (LLM) outputs to BLS data on gender representation in occupations. Each panel plots BLS % women (x-axis) against LLM % women (y-axis). The dotted gray diagonal indicates perfect parity (LLM = BLS). The five bias patterns are: Reality (accurate representation), Stereotype Exaggeration (amplified extremes, slope $> 1$ in logit space), Stereotype Reduction (dampened extremes, slope $< 1$ in logit space), Overrepresentation (uniform upward shift), and Underrepresentation (uniform downward shift). These schematic curves are illustrative rather than empirical, and are shown separately for clarity; in practice, models may exhibit combinations (e.g., overall underrepresentation combined with stereotype exaggeration).
  • Figure 2: Differences between four large language models’ representations of gender in occupational prompts compared to BLS benchmarks. Occupations are shown on the y-axis, with an “Average” row at the top summarizing mean differences across all 41 occupations. The x-axis shows percentage-point (pp) differences from BLS (absolute, not relative): negative values indicate underrepresentation and positive values indicate overrepresentation, with the dashed vertical line marking parity (0). Occupations are ordered from those where women are most overrepresented to most underrepresented. Only women are shown, with men serving as the implicit complement. Each dot represents one occupation, and points that would otherwise overlap are jittered vertically for readability.
  • Figure 3: Differences between four large language models’ representations of racial groups in occupational prompts compared to BLS benchmarks. Occupations are shown on the y-axis, with an “Average” row at the top summarizing mean differences across all 41 occupations. The x-axis shows percentage-point differences from BLS (absolute, not relative): negative values indicate underrepresentation and positive values indicate overrepresentation, with the dashed vertical line marking parity (0). Occupations are ordered from those where White workers are most overrepresented to those where they are most underrepresented. Separate panels show results for White, Hispanic, Black, and Asian workers. Each dot represents one model (ChatGPT = circle, Gemini = square, DeepSeek = diamond, Mistral = triangle); points that would otherwise overlap are jittered vertically for readability.
  • Figure 4: Systematic gender representation pooled across all models. Each point is an occupation, plotting BLS percentage of women against pooled LLM output. The diagonal line indicates parity. The fitted curve follows an S-shaped pattern, showing that models tend to exaggerate gender skews at both male- and female-dominated ends of the distribution.
  • Figure 5: Systematic racial representation pooled across all models. Each point is an occupation, plotting BLS percentages of racial groups against pooled LLM outputs. The diagonal line indicates parity. The fitted curves show underrepresentation of Black and White workers, overrepresentation of Hispanic and Asian workers, and stereotype exaggeration of demographic skews.
  • ...and 2 more figures