Camera Perspective Transformation to Bird's Eye View via Spatial Transformer Model for Road Intersection Monitoring

Rukesh Prajapati; Amr S. El-Wakeel

Camera Perspective Transformation to Bird's Eye View via Spatial Transformer Model for Road Intersection Monitoring

Rukesh Prajapati, Amr S. El-Wakeel

TL;DR

This work tackles the practical challenge of deriving a Bird's Eye View (BEV) of road intersections from a single camera to bridge the gap between BEV-rich simulators and real-world deployment. It introduces the Spatial-Transformer Double Decoder-UNet (SDD-UNet), a monocular BEV transformation model that uses a single encoder and dual decoders, with spatial-transformer-enabled skip connections to produce BEV masks while preserving vehicle localization. The approach is trained and evaluated in a simulated environment (CARLA RoadRunner via OpenStreetMap data) and demonstrates state-of-the-art performance with a Dice Similarity Coefficient around 0.957 and a centroid error of about 0.14 meters, outperforming a standard UNet and a ST-skip variant. The results indicate practical viability for real-time, drone-free BEV extraction from fixed cameras, enabling integration with BEV-based simulation-trained models for real-world road intersection monitoring and control.

Abstract

Road intersection monitoring and control research often utilize bird's eye view (BEV) simulators. In real traffic settings, achieving a BEV akin to that in a simulator necessitates the deployment of drones or specific sensor mounting, which is neither feasible nor practical. Consequently, traffic intersection management remains confined to simulation environments given these constraints. In this paper, we address the gap between simulated environments and real-world implementation by introducing a novel deep-learning model that converts a single camera's perspective of a road intersection into a BEV. We created a simulation environment that closely resembles a real-world traffic junction. The proposed model transforms the vehicles into BEV images, facilitating road intersection monitoring and control model processing. Inspired by image transformation techniques, we propose a Spatial-Transformer Double Decoder-UNet (SDD-UNet) model that aims to eliminate the transformed image distortions. In addition, the model accurately estimates the vehicle's positions and enables the direct application of simulation-trained models in real-world contexts. SDD-UNet model achieves an average dice similarity coefficient (DSC) above 95% which is 40% better than the original UNet model. The mean absolute error (MAE) is 0.102 and the centroid of the predicted mask is 0.14 meters displaced, on average, indicating high accuracy.

Camera Perspective Transformation to Bird's Eye View via Spatial Transformer Model for Road Intersection Monitoring

TL;DR

Abstract

Paper Structure (17 sections, 5 equations, 7 figures, 1 table)

This paper contains 17 sections, 5 equations, 7 figures, 1 table.

Introduction
Framework and Methodology
Simulation Environment and Dataset
Spatial-Transformer Double Decoder-UNet Model Architecture
Spatial Transformers
Localization Network
Grid Generator
Grid Sampler
Evaluation Metrics
Results and Discussion
Computer Vision Algorithmic Approach
Deep Learning Approach
Training Comparison
Comprehensive Results
Visual Analysis
...and 2 more sections

Figures (7)

Figure 1: Data collection process
Figure 2: Junction at West Virginia University replicated to create a simulation environment for data collection googlemaps
Figure 3: Architecture of proposed SDD-UNet model with single encoder and double decoder branch, where the first branch predicts a mask (op1) that comprises a bounding box for the vehicle inside the image and the second branch produces the mask (op2) for the BEV of vehicles position
Figure 4: Illustration of Spatial Transformer
Figure 5: Illustration of distortion in image from perspective transformation
...and 2 more figures

Camera Perspective Transformation to Bird's Eye View via Spatial Transformer Model for Road Intersection Monitoring

TL;DR

Abstract

Camera Perspective Transformation to Bird's Eye View via Spatial Transformer Model for Road Intersection Monitoring

Authors

TL;DR

Abstract

Table of Contents

Figures (7)