Table of Contents
Fetching ...

From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models

Zefan Cai, Haoyi Qiu, Haozhe Zhao, Ke Wan, Jiachen Li, Jiuxiang Gu, Wen Xiao, Nanyun Peng, Junjie Hu

TL;DR

This work investigates how alignment tuning with human-preference reward models shapes social bias in video diffusion models. It introduces VideoBiasEval, an event-centric, multi-granular framework that disentangles actor identity from content and tracks bias evolution from human preference data through reward models to aligned video outputs, including temporal stability metrics. The study finds that reward-alignment often amplifies existing gender and ethnicity biases and can stabilize these biases over time, though counter-bias reward models and controllable data curation can mitigate some disparities. The findings underscore the importance of bias-aware evaluation and propose a data-driven bias control direction—through controllable reward modeling and dataset composition—for more equitable video generation.

Abstract

Recent advances in video diffusion models have significantly enhanced text-to-video generation, particularly through alignment tuning using reward models trained on human preferences. While these methods improve visual quality, they can unintentionally encode and amplify social biases. To systematically trace how such biases evolve throughout the alignment pipeline, we introduce VideoBiasEval, a comprehensive diagnostic framework for evaluating social representation in video generation. Grounded in established social bias taxonomies, VideoBiasEval employs an event-based prompting strategy to disentangle semantic content (actions and contexts) from actor attributes (gender and ethnicity). It further introduces multi-granular metrics to evaluate (1) overall ethnicity bias, (2) gender bias conditioned on ethnicity, (3) distributional shifts in social attributes across model variants, and (4) the temporal persistence of bias within videos. Using this framework, we conduct the first end-to-end analysis connecting biases in human preference datasets, their amplification in reward models, and their propagation through alignment-tuned video diffusion models. Our results reveal that alignment tuning not only strengthens representational biases but also makes them temporally stable, producing smoother yet more stereotyped portrayals. These findings highlight the need for bias-aware evaluation and mitigation throughout the alignment process to ensure fair and socially responsible video generation.

From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models

TL;DR

This work investigates how alignment tuning with human-preference reward models shapes social bias in video diffusion models. It introduces VideoBiasEval, an event-centric, multi-granular framework that disentangles actor identity from content and tracks bias evolution from human preference data through reward models to aligned video outputs, including temporal stability metrics. The study finds that reward-alignment often amplifies existing gender and ethnicity biases and can stabilize these biases over time, though counter-bias reward models and controllable data curation can mitigate some disparities. The findings underscore the importance of bias-aware evaluation and propose a data-driven bias control direction—through controllable reward modeling and dataset composition—for more equitable video generation.

Abstract

Recent advances in video diffusion models have significantly enhanced text-to-video generation, particularly through alignment tuning using reward models trained on human preferences. While these methods improve visual quality, they can unintentionally encode and amplify social biases. To systematically trace how such biases evolve throughout the alignment pipeline, we introduce VideoBiasEval, a comprehensive diagnostic framework for evaluating social representation in video generation. Grounded in established social bias taxonomies, VideoBiasEval employs an event-based prompting strategy to disentangle semantic content (actions and contexts) from actor attributes (gender and ethnicity). It further introduces multi-granular metrics to evaluate (1) overall ethnicity bias, (2) gender bias conditioned on ethnicity, (3) distributional shifts in social attributes across model variants, and (4) the temporal persistence of bias within videos. Using this framework, we conduct the first end-to-end analysis connecting biases in human preference datasets, their amplification in reward models, and their propagation through alignment-tuned video diffusion models. Our results reveal that alignment tuning not only strengthens representational biases but also makes them temporally stable, producing smoother yet more stereotyped portrayals. These findings highlight the need for bias-aware evaluation and mitigation throughout the alignment process to ensure fair and socially responsible video generation.
Paper Structure (25 sections, 42 figures, 9 tables)

This paper contains 25 sections, 42 figures, 9 tables.

Figures (42)

  • Figure 1: Overview of our work: (1) We introduce VideoBiasEval, a bias evaluation framework for video generation that leverages event-based prompts and multi-granular metrics to assess ethnicity and gender bias (bottom left, \ref{['sec:evaluation_framework']}). The framework represents videos through social attribute annotations (top), where visual-language models (VLMs) infer actor attributes such as gender and ethnicity across frames and aggregate them for bias quantification \ref{['sec:evaluation_metrics']}). (2) We conduct the first comprehensive analysis of how image-based reward models, shaped by human-labeled preferences, influence the distribution of social attributes in diffusion-generated videos, disentangling human preference biases, reward preference biases, and their downstream impact on video diffusion model generation biases (bottom right, \ref{['sec:reward_datasets']}, \ref{['sec:reward_models']}, and \ref{['sec:alignment']}).
  • Figure 2: Illustration of videos generated by different diffusion models using varied prompt templates that specify actor attributes as detailed in \ref{['sec:event_prompting_template']}. The main character's social attributes, including gender and ethnicity, are extracted using our proposed VLM-based evaluation method described in \ref{['sec:evaluation_metrics']}.
  • Figure 3: Image examples of our constructed benchmark for evaluating preference in image reward models with generation prompts: "A/An [ethnicity][gender] is baking [context]." We only show the images with [gender]$\in \{$man, woman$\}$.
  • Figure 4: Action-level impact of alignment tuning guided by RM$_\text{M}$ and RM$_\text{W}$.
  • Figure 5: Ethnicity-aware gender bias (White).
  • ...and 37 more figures