Table of Contents
Fetching ...

BATIS: Bayesian Approaches for Targeted Improvement of Species Distribution Models

Catherine Villeneuve, Benjamin Akera, Mélisande Teng, David Rolnick

TL;DR

BATIS introduces a Bayesian framework that treats ML-derived SDM predictions as priors and iteratively updates them with limited ground observations to improve encounter-rate estimates across data-scarce regions. By explicitly modeling both aleatoric and epistemic uncertainty through Beta-Binomial updates and a suite of distributional uncertainty methods (e.g., MVN, HetReg), BATIS demonstrates rapid improvements in SDM reliability using minimal new data, validated on a large eBird-based benchmark with remote-sensing and WorldClim covariates. The study shows that uncertainty-aware updates outperform uncertainty-agnostic baselines, with aleatoric-focused approaches often yielding the strongest gains in low-data regimes, while highlighting ongoing challenges from data biases and scaling to large ecological patterns. These results indicate that BATIS offers a practical, computationally-light pathway to more trustworthy SDMs for conservation planning and resource allocation, with clear directions for incorporating ongoing citizen-science data streams.

Abstract

Species distribution models (SDMs), which aim to predict species occurrence based on environmental variables, are widely used to monitor and respond to biodiversity change. Recent deep learning advances for SDMs have been shown to perform well on complex and heterogeneous datasets, but their effectiveness remains limited by spatial biases in the data. In this paper, we revisit deep SDMs from a Bayesian perspective and introduce BATIS, a novel and practical framework wherein prior predictions are updated iteratively using limited observational data. Models must appropriately capture both aleatoric and epistemic uncertainty to effectively combine fine-grained local insights with broader ecological patterns. We benchmark an extensive set of uncertainty quantification approaches on a novel dataset including citizen science observations from the eBird platform. Our empirical study shows how Bayesian deep learning approaches can greatly improve the reliability of SDMs in data-scarce locations, which can contribute to ecological understanding and conservation efforts.

BATIS: Bayesian Approaches for Targeted Improvement of Species Distribution Models

TL;DR

BATIS introduces a Bayesian framework that treats ML-derived SDM predictions as priors and iteratively updates them with limited ground observations to improve encounter-rate estimates across data-scarce regions. By explicitly modeling both aleatoric and epistemic uncertainty through Beta-Binomial updates and a suite of distributional uncertainty methods (e.g., MVN, HetReg), BATIS demonstrates rapid improvements in SDM reliability using minimal new data, validated on a large eBird-based benchmark with remote-sensing and WorldClim covariates. The study shows that uncertainty-aware updates outperform uncertainty-agnostic baselines, with aleatoric-focused approaches often yielding the strongest gains in low-data regimes, while highlighting ongoing challenges from data biases and scaling to large ecological patterns. These results indicate that BATIS offers a practical, computationally-light pathway to more trustworthy SDMs for conservation planning and resource allocation, with clear directions for incorporating ongoing citizen-science data streams.

Abstract

Species distribution models (SDMs), which aim to predict species occurrence based on environmental variables, are widely used to monitor and respond to biodiversity change. Recent deep learning advances for SDMs have been shown to perform well on complex and heterogeneous datasets, but their effectiveness remains limited by spatial biases in the data. In this paper, we revisit deep SDMs from a Bayesian perspective and introduce BATIS, a novel and practical framework wherein prior predictions are updated iteratively using limited observational data. Models must appropriately capture both aleatoric and epistemic uncertainty to effectively combine fine-grained local insights with broader ecological patterns. We benchmark an extensive set of uncertainty quantification approaches on a novel dataset including citizen science observations from the eBird platform. Our empirical study shows how Bayesian deep learning approaches can greatly improve the reliability of SDMs in data-scarce locations, which can contribute to ecological understanding and conservation efforts.
Paper Structure (90 sections, 12 equations, 11 figures, 11 tables, 2 algorithms)

This paper contains 90 sections, 12 equations, 11 figures, 11 tables, 2 algorithms.

Figures (11)

  • Figure 1: Distribution of citizen science sampling trips across New Mexico, US (left) and South Africa (right), retrieved from the eBird database for the whole year of 2024. The value assigned to each 5km2 (US) and 10km2 (South Africa) grid cell corresponds to the total number of visits that were registered within that cell.
  • Figure 2: Iterative improvements for the different uncertainty estimation approaches with increasing number of checklist updates for the MAE, MSE and Top-10 metrics on the South Africa Region test set. We report the mean on three seeds and standard deviations for each model.
  • Figure 3: Evolution of the MAE in relation to the number of checklists (1, 6, 10) used to update the posterior distribution for three bird species of Kenya (Columba guinea, Ardea cinerea, Ploceus baglafecht) . The value assigned to each 70km2 grid cell corresponds to the mean MAE computed on the aggregation of all the hotspots located within that cell.
  • Figure 4: From left to right, for a) Training and b) Test : Geographical distribution of the total number of hotspots (left) and non-zero encounter rates (middle) associated with our Kenya subdataset, and distribution of the number of species encountered per hotspot (right).
  • Figure 5: From top to bottom, for a) Training and b) Test : Geographical distribution of the total number of hotspots (top) and non-zero encounter rates (middle) associated with our US-Winter subdataset, and distribution of the number of species encountered per hotspot (bottom).
  • ...and 6 more figures