OmniNWM: Omniscient Driving Navigation World Models
Bohan Li, Zhuang Ma, Dalong Du, Baorui Peng, Zhujin Liang, Zhenqiang Liu, Chao Ma, Yueming Jin, Hao Zhao, Wenjun Zeng, Xin Jin
TL;DR
OmniNWM addresses the need for unified, long-horizon driving world models by jointly forecasting panoramic RGB, semantics, depth, and 3D occupancy while enabling precise, pixel-level action control through normalized Plücker ray-maps. It introduces occupancy-grounded dense rewards to support closed-loop evaluation and planning, and a flexible forcing strategy to maintain generation quality over extended horizons, achieving state-of-the-art results in generation fidelity and control accuracy. The framework is complemented by OmniNWM-VLA and Tri-MIDI for planning, plus a Panoramic Diffusion Transformer that integrates multi-modal signals. Together, these elements yield strong zero-shot generalization and robust, occupancy-informed evaluation, highlighting the practical potential for holistic driving system development and testing.
Abstract
Autonomous driving world models are expected to work effectively across three core dimensions: state, action, and reward. Existing models, however, are typically restricted to limited state modalities, short video sequences, imprecise action control, and a lack of reward awareness. In this paper, we introduce OmniNWM, an omniscient panoramic navigation world model that addresses all three dimensions within a unified framework. For state, OmniNWM jointly generates panoramic videos of RGB, semantics, metric depth, and 3D occupancy. A flexible forcing strategy enables high-quality long-horizon auto-regressive generation. For action, we introduce a normalized panoramic Plucker ray-map representation that encodes input trajectories into pixel-level signals, enabling highly precise and generalizable control over panoramic video generation. Regarding reward, we move beyond learning reward functions with external image-based models: instead, we leverage the generated 3D occupancy to directly define rule-based dense rewards for driving compliance and safety. Extensive experiments demonstrate that OmniNWM achieves state-of-the-art performance in video generation, control accuracy, and long-horizon stability, while providing a reliable closed-loop evaluation framework through occupancy-grounded rewards. Project page is available at https://arlo0o.github.io/OmniNWM/.
