Table of Contents
Fetching ...

FlyAwareV2: A Multimodal Cross-Domain UAV Dataset for Urban Scene Understanding

Francesco Barbato, Matteo Caligiuri, Pietro Zanuttigh

TL;DR

FlyAwareV2 addresses the scarcity of large, annotated UAV datasets by offering a mixed-reality multimodal collection (RGB, depth, semantic labels) for urban-scene understanding under varied weather and lighting. The dataset integrates synthetic CARLA-based scenes with real imagery, augments depth for real data via Marigold, and provides extensive benchmarks for RGB and multimodal semantic segmentation, plus synthetic-to-real domain adaptation analyses. Key contributions include depth-enabled real data, diverse adverse-weather scenarios, and rigorous cross-domain experiments demonstrating synthetic data can train effective UAV perception models and that multimodal fusion enhances segmentation. The work enables robust, cross-domain UAV scene understanding with practical implications for autonomous navigation and safety in urban environments.

Abstract

The development of computer vision algorithms for Unmanned Aerial Vehicle (UAV) applications in urban environments heavily relies on the availability of large-scale datasets with accurate annotations. However, collecting and annotating real-world UAV data is extremely challenging and costly. To address this limitation, we present FlyAwareV2, a novel multimodal dataset encompassing both real and synthetic UAV imagery tailored for urban scene understanding tasks. Building upon the recently introduced SynDrone and FlyAware datasets, FlyAwareV2 introduces several new key contributions: 1) Multimodal data (RGB, depth, semantic labels) across diverse environmental conditions including varying weather and daytime; 2) Depth maps for real samples computed via state-of-the-art monocular depth estimation; 3) Benchmarks for RGB and multimodal semantic segmentation on standard architectures; 4) Studies on synthetic-to-real domain adaptation to assess the generalization capabilities of models trained on the synthetic data. With its rich set of annotations and environmental diversity, FlyAwareV2 provides a valuable resource for research on UAV-based 3D urban scene understanding.

FlyAwareV2: A Multimodal Cross-Domain UAV Dataset for Urban Scene Understanding

TL;DR

FlyAwareV2 addresses the scarcity of large, annotated UAV datasets by offering a mixed-reality multimodal collection (RGB, depth, semantic labels) for urban-scene understanding under varied weather and lighting. The dataset integrates synthetic CARLA-based scenes with real imagery, augments depth for real data via Marigold, and provides extensive benchmarks for RGB and multimodal semantic segmentation, plus synthetic-to-real domain adaptation analyses. Key contributions include depth-enabled real data, diverse adverse-weather scenarios, and rigorous cross-domain experiments demonstrating synthetic data can train effective UAV perception models and that multimodal fusion enhances segmentation. The work enables robust, cross-domain UAV scene understanding with practical implications for autonomous navigation and safety in urban environments.

Abstract

The development of computer vision algorithms for Unmanned Aerial Vehicle (UAV) applications in urban environments heavily relies on the availability of large-scale datasets with accurate annotations. However, collecting and annotating real-world UAV data is extremely challenging and costly. To address this limitation, we present FlyAwareV2, a novel multimodal dataset encompassing both real and synthetic UAV imagery tailored for urban scene understanding tasks. Building upon the recently introduced SynDrone and FlyAware datasets, FlyAwareV2 introduces several new key contributions: 1) Multimodal data (RGB, depth, semantic labels) across diverse environmental conditions including varying weather and daytime; 2) Depth maps for real samples computed via state-of-the-art monocular depth estimation; 3) Benchmarks for RGB and multimodal semantic segmentation on standard architectures; 4) Studies on synthetic-to-real domain adaptation to assess the generalization capabilities of models trained on the synthetic data. With its rich set of annotations and environmental diversity, FlyAwareV2 provides a valuable resource for research on UAV-based 3D urban scene understanding.
Paper Structure (18 sections, 7 figures, 10 tables)

This paper contains 18 sections, 7 figures, 10 tables.

Figures (7)

  • Figure 1: We introduce FlyAwareV2, a mixed-reality multimodal dataset for UAV imagery. We provide synthetic and real samples in varying weather conditions, with ground-truth depth information as well as semantic segmentation labels.
  • Figure 2: The FlyAwareV2 dataset provides color information, depth data and semantic labels for each frame.
  • Figure 3: The FlyAwareV2 dataset provides data in variable weather and daytime conditions.
  • Figure 4: Qualitative results: model trained on synthetic samples and tested on synthetic samples with varying weather conditions. First row: Input, Second row: Model Prediction, Third Row: Ground Truth.
  • Figure 5: Qualitative experiments: model trained on synthetic samples and tested on real samples with varying weather conditions. First row: Input, Second row: Model Prediction (no adaptation), Third Row: Model Prediction (UDA), Fourth Row: Ground Truth.
  • ...and 2 more figures