DC-Scene: Data-Centric Learning for 3D Scene Understanding

Ting Huang; Zeyu Zhang; Ruicheng Zhang; Yang Zhao

DC-Scene: Data-Centric Learning for 3D Scene Understanding

Ting Huang, Zeyu Zhang, Ruicheng Zhang, Yang Zhao

TL;DR

DC-Scene tackles the efficiency and data-scarcity hurdles in 3D scene understanding by introducing a CLIP-driven data quality filter and a dual-indicator quality mechanism coupled with a progressive curriculum. By selecting high-quality, semantically aligned scene-caption pairs and gradually incorporating more challenging samples, it achieves state-of-the-art CIDEr scores on ScanRefer and Nr3D while requiring roughly one-third of the training epochs. The approach is validated on multiple backbones (e.g., 3D CoCa, Vote2Cap-DETR++) and across two prominent datasets, demonstrating substantial training-time reductions without performance loss. This data-centric paradigm offers a generalizable path to more efficient 3D vision-language learning, with practical implications for robotics, AR/VR, and autonomous systems.

Abstract

3D scene understanding plays a fundamental role in vision applications such as robotics, autonomous driving, and augmented reality. However, advancing learning-based 3D scene understanding remains challenging due to two key limitations: (1) the large scale and complexity of 3D scenes lead to higher computational costs and slower training compared to 2D counterparts; and (2) high-quality annotated 3D datasets are significantly scarcer than those available for 2D vision. These challenges underscore the need for more efficient learning paradigms. In this work, we propose DC-Scene, a data-centric framework tailored for 3D scene understanding, which emphasizes enhancing data quality and training efficiency. Specifically, we introduce a CLIP-driven dual-indicator quality (DIQ) filter, combining vision-language alignment scores with caption-loss perplexity, along with a curriculum scheduler that progressively expands the training pool from the top 25% to 75% of scene-caption pairs. This strategy filters out noisy samples and significantly reduces dependence on large-scale labeled 3D data. Extensive experiments on ScanRefer and Nr3D demonstrate that DC-Scene achieves state-of-the-art performance (86.1 CIDEr with the top-75% subset vs. 85.4 with the full dataset) while reducing training cost by approximately two-thirds, confirming that a compact set of high-quality samples can outperform exhaustive training. Code will be available at https://github.com/AIGeeksGroup/DC-Scene.

DC-Scene: Data-Centric Learning for 3D Scene Understanding

TL;DR

Abstract

DC-Scene: Data-Centric Learning for 3D Scene Understanding

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (5)