Underwater Visual-Inertial-Acoustic-Depth SLAM with DVL Preintegration for Degraded Environments
Shuoshuo Ding, Tiedong Zhang, Dapeng Jiang, Ming Lei
TL;DR
Underwater SLAM often fails in degraded visibility due to limited visual features. This paper presents a graph-based framework that tightly fuses stereo camera, IMU, DVL, and pressure data, introducing a velocity-bias-based DVL preintegration and a hybrid visual tracking frontend to maintain robustness during visual degradation. The approach integrates multi-sensor residuals (visual, inertial, DVL, and depth) in a tight pose-graph optimization, enabling accurate localization and stable operation in challenging underwater environments. Experimental results in both simulated and real-world underwater settings show clear improvements in translation and rotation accuracy over state-of-the-art stereo VI-SLAM systems, with notable resilience when visual information is sparse or unreliable.
Abstract
Visual degradation caused by limited visibility, insufficient lighting, and feature scarcity in underwater environments presents significant challenges to visual-inertial simultaneous localization and mapping (SLAM) systems. To address these challenges, this paper proposes a graph-based visual-inertial-acoustic-depth SLAM system that integrates a stereo camera, an inertial measurement unit (IMU), the Doppler velocity log (DVL), and a pressure sensor. The key innovation lies in the tight integration of four distinct sensor modalities to ensure reliable operation, even under degraded visual conditions. To mitigate DVL drift and improve measurement efficiency, we propose a novel velocity-bias-based DVL preintegration strategy. At the frontend, hybrid tracking strategies and acoustic-inertial-depth joint optimization enhance system stability. Additionally, multi-source hybrid residuals are incorporated into a graph optimization framework. Extensive quantitative and qualitative analyses of the proposed system are conducted in both simulated and real-world underwater scenarios. The results demonstrate that our approach outperforms current state-of-the-art stereo visual-inertial SLAM systems in both stability and localization accuracy, exhibiting exceptional robustness, particularly in visually challenging environments.
