Improving the Physics of Video Generation with VJEPA-2 Reward Signal
Jianhao Yuan, Xiaofeng Zhang, Felix Friedrich, Nicolas Beltran-Velez, Melissa Hall, Reyhane Askari-Hemmat, Xiaochuang Han, Nicolas Ballas, Michal Drozdzal, Adriana Romero-Soriano
TL;DR
The paper tackles the gap between visual realism and physics understanding in video generation by coupling a state-of-the-art autoregressive video generator MAGI-1 with the self-supervised video model VJEPA-2. It introduces a surprise-based reward signal from VJEPA-2 to steer MAGI-1 during generation, augmenting the existing score with a penalty term and employing Best-of-N sampling to select low-surprise outputs. The approach achieves state-of-the-art physics plausibility on the PhysicsIQ benchmark, with roughly 6% gains over prior MAGI-1 baselines in both V2V and I2V generation. This work demonstrates that SSL-derived world models can guide physics-aware video synthesis, improving the plausibility of dynamic interactions in generated content and highlighting a promising direction for physics-constrained video generation.
Abstract
This is a short technical report describing the winning entry of the PhysicsIQ Challenge, presented at the Perception Test Workshop at ICCV 2025. State-of-the-art video generative models exhibit severely limited physical understanding, and often produce implausible videos. The Physics IQ benchmark has shown that visual realism does not imply physics understanding. Yet, intuitive physics understanding has shown to emerge from SSL pretraining on natural videos. In this report, we investigate whether we can leverage SSL-based video world models to improve the physics plausibility of video generative models. In particular, we build ontop of the state-of-the-art video generative model MAGI-1 and couple it with the recently introduced Video Joint Embedding Predictive Architecture 2 (VJEPA-2) to guide the generation process. We show that by leveraging VJEPA-2 as reward signal, we can improve the physics plausibility of state-of-the-art video generative models by ~6%.
