Table of Contents
Fetching ...

Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models

Xinmiao Huang, Qisong He, Zhenglin Huang, Boxuan Wang, Zhuoyun Li, Guangliang Cheng, Yi Dong, Xiaowei Huang

TL;DR

This work introduces Spatial-DISE, a unified, cognitively grounded benchmark for evaluating spatial reasoning in Vision-Language Models, anchored by a $2\times 2$ taxonomy that spans Intrinsic/Extrinsic and Static/Dynamic reasoning. It provides Spatial-DISE Bench (559 VQA pairs) and Spatial-DISE-12K (over 12K training pairs) generated via a Blender-based, scalable pipeline with three-stage curation to ensure quality and verifiability. Across 28 state-of-the-art models, results reveal a substantial gap to human performance, with multi-step dynamic tasks proving especially challenging and errors dominated by reasoning failures rather than perception. The framework enables precise diagnostic insights and paves the way for developing human-like spatial intelligence in VLMs, emphasizing robust generalization, interactive evaluation, and process-oriented reasoning outputs.

Abstract

Spatial reasoning ability is crucial for Vision Language Models (VLMs) to support real-world applications in diverse domains including robotics, augmented reality, and autonomous navigation. Unfortunately, existing benchmarks are inadequate in assessing spatial reasoning ability, especially the \emph{intrinsic-dynamic} spatial reasoning which is a fundamental aspect of human spatial cognition. In this paper, we propose a unified benchmark, \textbf{Spatial-DISE}, based on a cognitively grounded taxonomy that categorizes tasks into four fundamental quadrants: \textbf{I}ntrinsic-\textbf{S}tatic, Intrinsic-\textbf{D}ynamic, \textbf{E}xtrinsic-Static, and Extrinsic-Dynamic spatial reasoning. Moreover, to address the issue of data scarcity, we develop a scalable and automated pipeline to generate diverse and verifiable spatial reasoning questions, resulting in a new \textbf{Spatial-DISE} dataset that includes Spatial-DISE Bench (559 evaluation VQA pairs) and Spatial-DISE-12K (12K+ training VQA pairs). Our comprehensive evaluation across 28 state-of-the-art VLMs reveals that, current VLMs have a large and consistent gap to human competence, especially on multi-step multi-view spatial reasoning. Spatial-DISE offers a robust framework, valuable dataset, and clear direction for future research toward human-like spatial intelligence. Benchmark, dataset, and code will be publicly released.

Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models

TL;DR

This work introduces Spatial-DISE, a unified, cognitively grounded benchmark for evaluating spatial reasoning in Vision-Language Models, anchored by a taxonomy that spans Intrinsic/Extrinsic and Static/Dynamic reasoning. It provides Spatial-DISE Bench (559 VQA pairs) and Spatial-DISE-12K (over 12K training pairs) generated via a Blender-based, scalable pipeline with three-stage curation to ensure quality and verifiability. Across 28 state-of-the-art models, results reveal a substantial gap to human performance, with multi-step dynamic tasks proving especially challenging and errors dominated by reasoning failures rather than perception. The framework enables precise diagnostic insights and paves the way for developing human-like spatial intelligence in VLMs, emphasizing robust generalization, interactive evaluation, and process-oriented reasoning outputs.

Abstract

Spatial reasoning ability is crucial for Vision Language Models (VLMs) to support real-world applications in diverse domains including robotics, augmented reality, and autonomous navigation. Unfortunately, existing benchmarks are inadequate in assessing spatial reasoning ability, especially the \emph{intrinsic-dynamic} spatial reasoning which is a fundamental aspect of human spatial cognition. In this paper, we propose a unified benchmark, \textbf{Spatial-DISE}, based on a cognitively grounded taxonomy that categorizes tasks into four fundamental quadrants: \textbf{I}ntrinsic-\textbf{S}tatic, Intrinsic-\textbf{D}ynamic, \textbf{E}xtrinsic-Static, and Extrinsic-Dynamic spatial reasoning. Moreover, to address the issue of data scarcity, we develop a scalable and automated pipeline to generate diverse and verifiable spatial reasoning questions, resulting in a new \textbf{Spatial-DISE} dataset that includes Spatial-DISE Bench (559 evaluation VQA pairs) and Spatial-DISE-12K (12K+ training VQA pairs). Our comprehensive evaluation across 28 state-of-the-art VLMs reveals that, current VLMs have a large and consistent gap to human competence, especially on multi-step multi-view spatial reasoning. Spatial-DISE offers a robust framework, valuable dataset, and clear direction for future research toward human-like spatial intelligence. Benchmark, dataset, and code will be publicly released.
Paper Structure (37 sections, 8 equations, 13 figures, 15 tables, 5 algorithms)

This paper contains 37 sections, 8 equations, 13 figures, 15 tables, 5 algorithms.

Figures (13)

  • Figure 1: A Comprehensive Overview of the Spatial-DISE Framework, Generation Pipeline, and Benchmark Statistics. a) Comparison of examples from existing benchmarks, which primarily test general static reasoning, with cognition intrinsic-dynamic tasks from our Spatial-DISE benchmark. b) introduces the core DISE taxonomy, showing the four quadrants of spatial reasoning and their distribution in the 559-pair evaluation bench. c) presents evaluation results, showing a significant gap between model and human performance. d) details the synthetic data generation pipeline implemented in Blender, and e) provides a statistical breakdown of the task categories within both the Spatial-DISE Bench and the Spatial-DISE-12K.
  • Figure 2: 10 Tasks in Spatial-DISE Bench. Orange shows the Intrinsic-Dynamic Tasks, Green shows the Intrinsic-Static Tasks, Pink shows the Extrinsic-Static Tasks and Blue shows the Extrinsic-Dynamic Tasks.
  • Figure 3: Synthetic Data Generation and Quality Control.
  • Figure 4: Error example of Failure in Rule Application.
  • Figure 5: Synthetic 3D Rotation Data Example.
  • ...and 8 more figures