Prompt-Guided Spatial Understanding with RGB-D Transformers for Fine-Grained Object Relation Reasoning

Tanner Muturi; Blessing Agyei Kyem; Joshua Kofi Asamoah; Neema Jakisa Owor; Richard Dyzinela; Andrews Danyo; Yaw Adu-Gyamfi; Armstrong Aboah

Prompt-Guided Spatial Understanding with RGB-D Transformers for Fine-Grained Object Relation Reasoning

Tanner Muturi, Blessing Agyei Kyem, Joshua Kofi Asamoah, Neema Jakisa Owor, Richard Dyzinela, Andrews Danyo, Yaw Adu-Gyamfi, Armstrong Aboah

TL;DR

The paper tackles robust spatial reasoning in cluttered warehouse scenes by augmenting RGB-D transformers with prompt-based geometric grounding. It introduces a SpatialBot-based architecture that embeds object bounding box coordinates into prompts and uses an answer normalization module to align outputs with evaluation protocols, trained on the Physical AI Spatial Intelligence Warehouse dataset across four spatial tasks. The method demonstrates that explicit spatial grounding and depth cues substantially improve performance, achieving an S1 score of 73.0606 and ranking 4th on Track 3 of the AI City Challenge. This work offers a practical path toward depth-aware, geometry-grounded vision-language systems for industrial applications, reducing reliance on 2D appearance cues and enabling reliable multi-object reasoning.

Abstract

Spatial reasoning in large-scale 3D environments such as warehouses remains a significant challenge for vision-language systems due to scene clutter, occlusions, and the need for precise spatial understanding. Existing models often struggle with generalization in such settings, as they rely heavily on local appearance and lack explicit spatial grounding. In this work, we introduce a dedicated spatial reasoning framework for the Physical AI Spatial Intelligence Warehouse dataset introduced in the Track 3 2025 AI City Challenge. Our approach enhances spatial comprehension by embedding mask dimensions in the form of bounding box coordinates directly into the input prompts, enabling the model to reason over object geometry and layout. We fine-tune the framework across four question categories namely: Distance Estimation, Object Counting, Multi-choice Grounding, and Spatial Relation Inference using task-specific supervision. To further improve consistency with the evaluation system, normalized answers are appended to the GPT response within the training set. Our comprehensive pipeline achieves a final score of 73.0606, placing 4th overall on the public leaderboard. These results demonstrate the effectiveness of structured prompt enrichment and targeted optimization in advancing spatial reasoning for real-world industrial environments.

Prompt-Guided Spatial Understanding with RGB-D Transformers for Fine-Grained Object Relation Reasoning

TL;DR

Abstract

Prompt-Guided Spatial Understanding with RGB-D Transformers for Fine-Grained Object Relation Reasoning

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (4)