Table of Contents
Fetching ...

DSI-Bench: A Benchmark for Dynamic Spatial Intelligence

Ziang Zhang, Zehan Wang, Guanghao Zhang, Weilong Dai, Yan Xia, Ziang Yan, Minjie Hong, Zhou Zhao

TL;DR

This work introduces Dynamic Spatial Intelligence and presents DSI-Bench, a benchmark for evaluating dynamic spatial reasoning in 3D scenes. It provides nearly 1,000 dynamic videos and over 1,700 manually annotated VQA pairs, with symmetry-based augmentation to mitigate motion biases, and categorizes tasks across Object–Scene, Observer–Scene, and Observer–Object relations. Comprehensive evaluation of 14 VLMs and 3D spatial-expertise models reveals that dynamic reasoning remains challenging, with VLMs showing bias and hallucination tendencies, limited robustness, and that larger models improve perceptual accuracy but not robustness. The results establish DSI-Bench as a resource for advancing models toward robust dynamic spatial perception and reasoning in real-world settings.

Abstract

Reasoning about dynamic spatial relationships is essential, as both observers and objects often move simultaneously. Although vision-language models (VLMs) and visual expertise models excel in 2D tasks and static scenarios, their ability to fully understand dynamic 3D scenarios remains limited. We introduce Dynamic Spatial Intelligence and propose DSI-Bench, a benchmark with nearly 1,000 dynamic videos and over 1,700 manually annotated questions covering nine decoupled motion patterns of observers and objects. Spatially and temporally symmetric designs reduce biases and enable systematic evaluation of models' reasoning about self-motion and object motion. Our evaluation of 14 VLMs and expert models reveals key limitations: models often conflate observer and object motion, exhibit semantic biases, and fail to accurately infer relative relationships in dynamic scenarios. Our DSI-Bench provides valuable findings and insights about the future development of general and expertise models with dynamic spatial intelligence.

DSI-Bench: A Benchmark for Dynamic Spatial Intelligence

TL;DR

This work introduces Dynamic Spatial Intelligence and presents DSI-Bench, a benchmark for evaluating dynamic spatial reasoning in 3D scenes. It provides nearly 1,000 dynamic videos and over 1,700 manually annotated VQA pairs, with symmetry-based augmentation to mitigate motion biases, and categorizes tasks across Object–Scene, Observer–Scene, and Observer–Object relations. Comprehensive evaluation of 14 VLMs and 3D spatial-expertise models reveals that dynamic reasoning remains challenging, with VLMs showing bias and hallucination tendencies, limited robustness, and that larger models improve perceptual accuracy but not robustness. The results establish DSI-Bench as a resource for advancing models toward robust dynamic spatial perception and reasoning in real-world settings.

Abstract

Reasoning about dynamic spatial relationships is essential, as both observers and objects often move simultaneously. Although vision-language models (VLMs) and visual expertise models excel in 2D tasks and static scenarios, their ability to fully understand dynamic 3D scenarios remains limited. We introduce Dynamic Spatial Intelligence and propose DSI-Bench, a benchmark with nearly 1,000 dynamic videos and over 1,700 manually annotated questions covering nine decoupled motion patterns of observers and objects. Spatially and temporally symmetric designs reduce biases and enable systematic evaluation of models' reasoning about self-motion and object motion. Our evaluation of 14 VLMs and expert models reveals key limitations: models often conflate observer and object motion, exhibit semantic biases, and fail to accurately infer relative relationships in dynamic scenarios. Our DSI-Bench provides valuable findings and insights about the future development of general and expertise models with dynamic spatial intelligence.
Paper Structure (38 sections, 12 figures, 3 tables)

This paper contains 38 sections, 12 figures, 3 tables.

Figures (12)

  • Figure 1: Dynamic Spatial Intelligence. Unlike static settings, dynamic scenarios involve evolving spatial relationships among the observer, observed objects, and the environment. Humans can intuitively perceive such changes in spatial relations, whereas Vision-Language Models (VLMs) often exhibit hallucinations and biases in dynamic spatial reasoning due to semantic misleadness and coupled motion understanding.
  • Figure 2: Left: Task distribution in DSI-Bench; Middle: Observer Motion distribution in DSI-Bench; Right: Observed Motion distribution in DSI-Bench
  • Figure 3: Illustration of the DSI-Bench construction pipeline: videos are sampled from diverse motion datasets, QA tasks and options are template-based genrated and manually refined, and videos are augmented with spatio-temporal flipping to mitigate data bias.
  • Figure 4: Proportion of models selected with the "forward" option and ground-truth annotations containing "forward".
  • Figure 5: Performance gap of VLMs between static and dynamic conditions.
  • ...and 7 more figures