RealDPO: Real or Not Real, that is the Preference
Guo Cheng, Danni Yang, Ziqi Huang, Jianlou Si, Chenyang Si, Ziwei Liu
TL;DR
This work tackles the difficulty of producing natural, contextually coherent human motion in video generation by proposing RealDPO, a real-data–driven Direct Preference Optimization workflow that avoids reward-model pitfalls. By treating diffusion-based video generation as a diffusion MDP and optimizing with win-lose preferences where real videos serve as wins, RealDPO achieves stronger visual-text alignment and motion realism with a data-efficient approach. It introduces RealAction-5K, a compact, high-quality dataset of daily actions to guide preference learning, and employs a tailored loss with an EMA-updated reference model to iteratively refine outputs. Empirical results across human studies and automated metrics show RealDPO outperforming SFT, LiFT, VideoAlign, and reward-model–based methods, underscoring its practical impact for action-centric video synthesis and its potential generalization to broader domains.
Abstract
Video generative models have recently achieved notable advancements in synthesis quality. However, generating complex motions remains a critical challenge, as existing models often struggle to produce natural, smooth, and contextually consistent movements. This gap between generated and real-world motions limits their practical applicability. To address this issue, we introduce RealDPO, a novel alignment paradigm that leverages real-world data as positive samples for preference learning, enabling more accurate motion synthesis. Unlike traditional supervised fine-tuning (SFT), which offers limited corrective feedback, RealDPO employs Direct Preference Optimization (DPO) with a tailored loss function to enhance motion realism. By contrasting real-world videos with erroneous model outputs, RealDPO enables iterative self-correction, progressively refining motion quality. To support post-training in complex motion synthesis, we propose RealAction-5K, a curated dataset of high-quality videos capturing human daily activities with rich and precise motion details. Extensive experiments demonstrate that RealDPO significantly improves video quality, text alignment, and motion realism compared to state-of-the-art models and existing preference optimization techniques.
