Progressive Multi-Modal Fusion for Robust 3D Object Detection

Rohit Mohan; Daniele Cattaneo; Florian Drews; Abhinav Valada

Progressive Multi-Modal Fusion for Robust 3D Object Detection

Rohit Mohan, Daniele Cattaneo, Florian Drews, Abhinav Valada

TL;DR

This work tackles robust 3D object detection in autonomous driving by introducing ProFusion3D, a progressive fusion framework that fuses LiDAR and multi-view camera features in both BEV and PV at intermediate feature and object-query levels. It combines an inter-intra fusion module with dual BEV/PV decoders and a joint decoder to leverage local and global context, while a self-supervised multi-modal mask modeling pre-training scheme enhances data efficiency and cross-modal representation learning. On nuScenes and Argoverse2, ProFusion3D achieves state-of-the-art performance and demonstrates strong robustness when one modality is unavailable. The approach also yields significant data-efficiency gains from pre-training and provides detailed ablations illustrating the contributions of fusion strategy, pre-training objectives, and decoder design, highlighting practical benefits for real-world autonomous driving perception.

Abstract

Multi-sensor fusion is crucial for accurate 3D object detection in autonomous driving, with cameras and LiDAR being the most commonly used sensors. However, existing methods perform sensor fusion in a single view by projecting features from both modalities either in Bird's Eye View (BEV) or Perspective View (PV), thus sacrificing complementary information such as height or geometric proportions. To address this limitation, we propose ProFusion3D, a progressive fusion framework that combines features in both BEV and PV at both intermediate and object query levels. Our architecture hierarchically fuses local and global features, enhancing the robustness of 3D object detection. Additionally, we introduce a self-supervised mask modeling pre-training strategy to improve multi-modal representation learning and data efficiency through three novel objectives. Extensive experiments on nuScenes and Argoverse2 datasets conclusively demonstrate the efficacy of ProFusion3D. Moreover, ProFusion3D is robust to sensor failure, demonstrating strong performance when only one modality is available.

Progressive Multi-Modal Fusion for Robust 3D Object Detection

TL;DR

Abstract

Progressive Multi-Modal Fusion for Robust 3D Object Detection

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (8)