Table of Contents
Fetching ...

Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning

Xiangyu Meng, Zixian Zhang, Zhenghao Zhang, Junchao Liao, Long Qin, Weizhi Wang

TL;DR

Identity-GRPO introduces a preference-driven reinforcement learning framework to tackle multi-human identity-preserving video generation. It builds a large-scale, hybridly labeled preference dataset and trains an identity-consistent reward model using both auto-labeled and human-labeled data with consistency filtering and a cosine-scheduled curriculum. The method then applies Group Relative Policy Optimization to MH-IPV, leveraging flow-based video sampling and stability strategies such as prompt finetuning and initial-noise differentiation. Empirical results show up to 18.9% improvements in identity-consistency metrics over VACE and Phantom, with favorable user-study outcomes, highlighting the effectiveness of aligning RL with human preferences for personalized, multi-human video generation.

Abstract

While advanced methods like VACE and Phantom have advanced video generation for specific subjects in diverse scenarios, they struggle with multi-human identity preservation in dynamic interactions, where consistent identities across multiple characters are critical. To address this, we propose Identity-GRPO, a human feedback-driven optimization pipeline for refining multi-human identity-preserving video generation. First, we construct a video reward model trained on a large-scale preference dataset containing human-annotated and synthetic distortion data, with pairwise annotations focused on maintaining human consistency throughout the video. We then employ a GRPO variant tailored for multi-human consistency, which greatly enhances both VACE and Phantom. Through extensive ablation studies, we evaluate the impact of annotation quality and design choices on policy optimization. Experiments show that Identity-GRPO achieves up to 18.9% improvement in human consistency metrics over baseline methods, offering actionable insights for aligning reinforcement learning with personalized video generation.

Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning

TL;DR

Identity-GRPO introduces a preference-driven reinforcement learning framework to tackle multi-human identity-preserving video generation. It builds a large-scale, hybridly labeled preference dataset and trains an identity-consistent reward model using both auto-labeled and human-labeled data with consistency filtering and a cosine-scheduled curriculum. The method then applies Group Relative Policy Optimization to MH-IPV, leveraging flow-based video sampling and stability strategies such as prompt finetuning and initial-noise differentiation. Empirical results show up to 18.9% improvements in identity-consistency metrics over VACE and Phantom, with favorable user-study outcomes, highlighting the effectiveness of aligning RL with human preferences for personalized, multi-human video generation.

Abstract

While advanced methods like VACE and Phantom have advanced video generation for specific subjects in diverse scenarios, they struggle with multi-human identity preservation in dynamic interactions, where consistent identities across multiple characters are critical. To address this, we propose Identity-GRPO, a human feedback-driven optimization pipeline for refining multi-human identity-preserving video generation. First, we construct a video reward model trained on a large-scale preference dataset containing human-annotated and synthetic distortion data, with pairwise annotations focused on maintaining human consistency throughout the video. We then employ a GRPO variant tailored for multi-human consistency, which greatly enhances both VACE and Phantom. Through extensive ablation studies, we evaluate the impact of annotation quality and design choices on policy optimization. Experiments show that Identity-GRPO achieves up to 18.9% improvement in human consistency metrics over baseline methods, offering actionable insights for aligning reinforcement learning with personalized video generation.
Paper Structure (21 sections, 10 equations, 2 figures, 4 tables)

This paper contains 21 sections, 10 equations, 2 figures, 4 tables.

Figures (2)

  • Figure 1: (a) and (b) respectively show the performance curves of Identity-GRPO on VACE-1.3B and Phantom-1.3B. Both exhibit a clear upward trend.
  • Figure 2: Visualization results for qualitative analysis. The first two groups show a comparison between VACE-1.3B and VACE-1.3B+Identity-GRPO, while the last two groups compare Phantom-1.3B with Phantom-1.3B+Identity-GRPO. In each group, the first row presents the results from the baseline model, and the second row shows the results generated by Identity-GRPO.