Table of Contents
Fetching ...

Freehand 3D Ultrasound Imaging: Sim-in-the-Loop Probe Pose Optimization via Visual Servoing

Yameng Zhang, Dianye Huang, Max Q. -H. Meng, Nassir Navab, Zhongliang Jiang

TL;DR

This work tackles the pose-estimation bottleneck in freehand 3D ultrasound by introducing a cost-effective, camera-based system that uses two eye-in-hand cameras observing a textured planar workspace. A simulation-in-the-loop framework couples real-world observations with a PBVS controller in a simulated environment (CoppeliaSim) to iteratively minimize pose error, augmented by an image restoration step to handle occlusions and lighting variations, and a one-time Sim2Real calibration to bridge simulation and reality. The approach is validated on a soft vascular phantom, a 3D-printed conical model, and a human arm, achieving sub-millimeter Hausdorff distances and robust translational accuracy across single and dual-camera configurations; it reduces reliance on costly tracking devices and improves stability over long freehand sweeps. The results demonstrate practical potential for low-cost, accurate freehand 3D US in clinical and training settings, with future work focused on real-time acceleration, broader probe types, and prospective clinical trials.

Abstract

Freehand 3D ultrasound (US) imaging using conventional 2D probes offers flexibility and accessibility for diverse clinical applications but faces challenges in accurate probe pose estimation. Traditional methods depend on costly tracking systems, while neural network-based methods struggle with image noise and error accumulation, compromising reconstruction precision. We propose a cost-effective and versatile solution that leverages lightweight cameras and visual servoing in simulated environments for precise 3D US imaging. These cameras capture visual feedback from a textured planar workspace. To counter occlusions and lighting issues, we introduce an image restoration method that reconstructs occluded regions by matching surrounding texture patterns. For pose estimation, we develop a simulation-in-the-loop approach, which replicates the system setup in simulation and iteratively minimizes pose errors between simulated and real-world observations. A visual servoing controller refines the alignment of camera views, improving translational estimation by optimizing image alignment. Validations on a soft vascular phantom, a 3D-printed conical model, and a human arm demonstrate the robustness and accuracy of our approach, with Hausdorff distances to the reference reconstructions of 0.359 mm, 1.171 mm, and 0.858 mm, respectively. These results confirm the method's potential for reliable freehand 3D US reconstruction.

Freehand 3D Ultrasound Imaging: Sim-in-the-Loop Probe Pose Optimization via Visual Servoing

TL;DR

This work tackles the pose-estimation bottleneck in freehand 3D ultrasound by introducing a cost-effective, camera-based system that uses two eye-in-hand cameras observing a textured planar workspace. A simulation-in-the-loop framework couples real-world observations with a PBVS controller in a simulated environment (CoppeliaSim) to iteratively minimize pose error, augmented by an image restoration step to handle occlusions and lighting variations, and a one-time Sim2Real calibration to bridge simulation and reality. The approach is validated on a soft vascular phantom, a 3D-printed conical model, and a human arm, achieving sub-millimeter Hausdorff distances and robust translational accuracy across single and dual-camera configurations; it reduces reliance on costly tracking devices and improves stability over long freehand sweeps. The results demonstrate practical potential for low-cost, accurate freehand 3D US in clinical and training settings, with future work focused on real-time acceleration, broader probe types, and prospective clinical trials.

Abstract

Freehand 3D ultrasound (US) imaging using conventional 2D probes offers flexibility and accessibility for diverse clinical applications but faces challenges in accurate probe pose estimation. Traditional methods depend on costly tracking systems, while neural network-based methods struggle with image noise and error accumulation, compromising reconstruction precision. We propose a cost-effective and versatile solution that leverages lightweight cameras and visual servoing in simulated environments for precise 3D US imaging. These cameras capture visual feedback from a textured planar workspace. To counter occlusions and lighting issues, we introduce an image restoration method that reconstructs occluded regions by matching surrounding texture patterns. For pose estimation, we develop a simulation-in-the-loop approach, which replicates the system setup in simulation and iteratively minimizes pose errors between simulated and real-world observations. A visual servoing controller refines the alignment of camera views, improving translational estimation by optimizing image alignment. Validations on a soft vascular phantom, a 3D-printed conical model, and a human arm demonstrate the robustness and accuracy of our approach, with Hausdorff distances to the reference reconstructions of 0.359 mm, 1.171 mm, and 0.858 mm, respectively. These results confirm the method's potential for reliable freehand 3D US reconstruction.
Paper Structure (20 sections, 10 equations, 10 figures, 3 tables, 1 algorithm)

This paper contains 20 sections, 10 equations, 10 figures, 3 tables, 1 algorithm.

Figures (10)

  • Figure 1: Illustration of a freehand 3D US imaging system using a conventional 2D US probe. The acquired 2D images can be stacked to reconstruct 3D volumes based on positional tracking data from tracking devices or pose estimation algorithms.
  • Figure 2: Sim-in-the-loop US probe pose estimation system: Two monocular cameras on the probe capture visual features from a planar RGB-patterned workspace. Real-world images are transmitted to the simulation, where visual servoing aligns them with simulated images, enabling extraction of the probe pose in the global frame.
  • Figure 3: Pipeline of the Probe Pose Estimation Algorithm. The process begins with image restoration, which recovers details obscured by objects or lighting reflections. Next, the pose error is computed by comparing the restored image with a simulated counterpart. A PBVS approach iteratively updates the probe's location in the simulation to achieve the best match with the restored real-world camera observation. By applying a one-time Sim2Real compensation to the optimized probe pose in simulation, we obtain a refined pose estimation. The robotic arm is included solely for quantitative evaluation, not as part of the proposed pose estimation method.
  • Figure 4: Sim2Real transformation chain. The calibration process estimates the offset transformations $\mathbf{T}_{\mathcal{V}}^{\mathcal{W}}$ and $\mathbf{T}_{\mathcal{P}}^{\mathcal{Q}}$ to align the simulated and real probe poses for accurate 3D US reconstruction.
  • Figure 5: Image restoration results: (a) image affected by lighting reflection, (b) image with arm occlusion, and (c) image captured in dim lighting. (d), (e), and (f) show the restored versions of (a), (b), and (c), respectively.
  • ...and 5 more figures