Table of Contents
Fetching ...

ReefNet: A Large scale, Taxonomically Enriched Dataset and Benchmark for Hard Coral Classification

Yahia Battach, Abdulwahab Felemban, Faizan Farooq Khan, Yousef A. Radwan, Xiang Li, Fabio Marchese, Sara Beery, Burton H. Jones, Francesca Benzoni, Mohamed Elhoseiny

TL;DR

ReefNet provides a large, WoRMS-aligned dataset of hard coral images with genus-level labels gathered from 76 CoralNet sources plus an Al-Wajh Red Sea subset, totaling approximately 925K annotations across 44 genera and 334K images. Two benchmarking settings—within-source and cross-source—are introduced to evaluate in-domain performance and domain generalization across diverse reef sites, including a dedicated Al-Wajh test. Experiments show strong within-site performance but pronounced degradation under domain shift, especially for rare or morphologically similar genera, while zero-shot approaches lag behind fine-tuned models though benefits from taxonomic pretraining and contextual descriptions. By releasing dataset, code, and pretrained models, ReefNet aims to catalyze robust, domain-adaptive coral monitoring and broader ecological insights while highlighting challenges in cross-region generalization and taxonomic fine-grained recognition.

Abstract

Coral reefs are rapidly declining due to anthropogenic pressures such as climate change, underscoring the urgent need for scalable, automated monitoring. We introduce ReefNet, a large public coral reef image dataset with point-label annotations mapped to the World Register of Marine Species (WoRMS). ReefNet aggregates imagery from 76 curated CoralNet sources and an additional site from Al Wajh in the Red Sea, totaling approximately 925000 genus-level hard coral annotations with expert-verified labels. Unlike prior datasets, which are often limited by size, geography, or coarse labels and are not ML-ready, ReefNet offers fine-grained, taxonomically mapped labels at a global scale to WoRMS. We propose two evaluation settings: (i) a within-source benchmark that partitions each source's images for localized evaluation, and (ii) a cross-source benchmark that withholds entire sources to test domain generalization. We analyze both supervised and zero-shot classification performance on ReefNet and find that while supervised within-source performance is promising, supervised performance drops sharply across domains, and performance is low across the board for zero-shot models, especially for rare and visually similar genera. This provides a challenging benchmark intended to catalyze advances in domain generalization and fine-grained coral classification. We will release our dataset, benchmarking code, and pretrained models to advance robust, domain-adaptive, global coral reef monitoring and conservation.

ReefNet: A Large scale, Taxonomically Enriched Dataset and Benchmark for Hard Coral Classification

TL;DR

ReefNet provides a large, WoRMS-aligned dataset of hard coral images with genus-level labels gathered from 76 CoralNet sources plus an Al-Wajh Red Sea subset, totaling approximately 925K annotations across 44 genera and 334K images. Two benchmarking settings—within-source and cross-source—are introduced to evaluate in-domain performance and domain generalization across diverse reef sites, including a dedicated Al-Wajh test. Experiments show strong within-site performance but pronounced degradation under domain shift, especially for rare or morphologically similar genera, while zero-shot approaches lag behind fine-tuned models though benefits from taxonomic pretraining and contextual descriptions. By releasing dataset, code, and pretrained models, ReefNet aims to catalyze robust, domain-adaptive coral monitoring and broader ecological insights while highlighting challenges in cross-region generalization and taxonomic fine-grained recognition.

Abstract

Coral reefs are rapidly declining due to anthropogenic pressures such as climate change, underscoring the urgent need for scalable, automated monitoring. We introduce ReefNet, a large public coral reef image dataset with point-label annotations mapped to the World Register of Marine Species (WoRMS). ReefNet aggregates imagery from 76 curated CoralNet sources and an additional site from Al Wajh in the Red Sea, totaling approximately 925000 genus-level hard coral annotations with expert-verified labels. Unlike prior datasets, which are often limited by size, geography, or coarse labels and are not ML-ready, ReefNet offers fine-grained, taxonomically mapped labels at a global scale to WoRMS. We propose two evaluation settings: (i) a within-source benchmark that partitions each source's images for localized evaluation, and (ii) a cross-source benchmark that withholds entire sources to test domain generalization. We analyze both supervised and zero-shot classification performance on ReefNet and find that while supervised within-source performance is promising, supervised performance drops sharply across domains, and performance is low across the board for zero-shot models, especially for rare and visually similar genera. This provides a challenging benchmark intended to catalyze advances in domain generalization and fine-grained coral classification. We will release our dataset, benchmarking code, and pretrained models to advance robust, domain-adaptive, global coral reef monitoring and conservation.
Paper Structure (50 sections, 16 figures, 12 tables)

This paper contains 50 sections, 16 figures, 12 tables.

Figures (16)

  • Figure 1: Overview of the Hierarchical Structure of ReefNet: a curated dataset of coral reef annotations focused on hard corals (Scleractinia), with textual descriptions for each genus.
  • Figure 2: Log-scale Distribution of Annotations Across Hard Coral Taxa. The plot includes 43 genera and one family-level class (Fungiidae).
  • Figure 3: Geographic Distribution of the Annotations. Sources located within two degrees of Latitude and Longitude are grouped into a single point. One source did not contain any location data and is displayed in Antarctica. Marine Ecoregions of the World are shown in light blue spalding_ecoregions.
  • Figure 4: Qualitative examples. Ground truth is shown above and model prediction below; the rightmost image illustrates the challenge of assigning a single label to patches with multiple corals.
  • Figure 5: ReefNet Curation Pipeline. The diagram illustrates the sequential filtering and exclusion steps applied to 1,366 CoralNet image sources to arrive at the final ReefNet dataset of 76 sources. Sources were excluded due to insufficient number of human-verified annotations, ecological or privacy concerns, and calibration-related issues.
  • ...and 11 more figures