DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
Ziang Zhang, Zehan Wang, Guanghao Zhang, Weilong Dai, Yan Xia, Ziang Yan, Minjie Hong, Zhou Zhao
TL;DR
This work introduces Dynamic Spatial Intelligence and presents DSI-Bench, a benchmark for evaluating dynamic spatial reasoning in 3D scenes. It provides nearly 1,000 dynamic videos and over 1,700 manually annotated VQA pairs, with symmetry-based augmentation to mitigate motion biases, and categorizes tasks across Object–Scene, Observer–Scene, and Observer–Object relations. Comprehensive evaluation of 14 VLMs and 3D spatial-expertise models reveals that dynamic reasoning remains challenging, with VLMs showing bias and hallucination tendencies, limited robustness, and that larger models improve perceptual accuracy but not robustness. The results establish DSI-Bench as a resource for advancing models toward robust dynamic spatial perception and reasoning in real-world settings.
Abstract
Reasoning about dynamic spatial relationships is essential, as both observers and objects often move simultaneously. Although vision-language models (VLMs) and visual expertise models excel in 2D tasks and static scenarios, their ability to fully understand dynamic 3D scenarios remains limited. We introduce Dynamic Spatial Intelligence and propose DSI-Bench, a benchmark with nearly 1,000 dynamic videos and over 1,700 manually annotated questions covering nine decoupled motion patterns of observers and objects. Spatially and temporally symmetric designs reduce biases and enable systematic evaluation of models' reasoning about self-motion and object motion. Our evaluation of 14 VLMs and expert models reveals key limitations: models often conflate observer and object motion, exhibit semantic biases, and fail to accurately infer relative relationships in dynamic scenarios. Our DSI-Bench provides valuable findings and insights about the future development of general and expertise models with dynamic spatial intelligence.
