Table of Contents
Fetching ...

Ultracool dwarf Science with MachIne LEarning (USMILE). I. Scalable Tree-Based Models for Photometric Spectral Classification and New Discoveries from LSST Data Preview 1 and Euclid Quick Data Release 1

Zhoujian Zhang, Yanxia Li

TL;DR

The paper tackles the challenge of discovering ultracool dwarfs in the data-rich LSST and Euclid era by building USMILE Avocado, a scalable two-model framework (classifier and regressor) based on gradient-boosted trees that can handle missing photometric data. It leverages a large, uncertainty-augmented labeled set derived from UltracoolSheet, reddened stars, and quasars, with eight colors anchored on $y_{ m LSST}$ computed from synthesized LSST, VHS, and CatWISE photometry, and validated against Euclid QDR1 spectra. The results show ROC AUC $=0.976$, F1 $=0.92$ for classification, and regressor MSE $=0.88$ subtypes, with external spectroscopic validation from Euclid confirming 15 new M6–L2 ultracool dwarfs and guiding reliability regimes, plus 25 additional candidates. The work demonstrates scalable, bias-resistant photometric spectral classification using ML in the LSST-Euclid era, enabling a more complete ultracool-dwarf census and informing follow-up strategies.

Abstract

We present the Ultracool dwarf Science with MachIne LEarning (USMILE), a program developing machine-learning tools for the discovery and characterization of ultracool dwarfs. We introduce USMILE Avocado, a spectral classification framework that uses broadband photometry from wide-field surveys -- Rubin Observatory LSST Data Preview 1, VISTA Hemisphere Survey, and CatWISE -- as input features. The framework has two gradient-boosted decision-tree models scalable to the massive data volumes of modern surveys: the classifier, which distinguishes ultracool dwarfs from stellar/extragalactic contaminants, and the regressor, which predicts spectral types. A key strength is its ability to natively handle missing photometric features, whereas earlier machine-learning approaches required complete multi-band detections or relied on imputation, thereby excluding genuine ultracool dwarfs or introducing bias. Trained on an augmented labeled dataset of >2 million sources built from known ultracool dwarfs, reddened early-type stars, and quasars, the models achieve strong performance: the classifier attains an ROC AUC of 0.976 and an F1 score of 0.92, while the regressor yields a mean-squared error of 0.88 subtypes. Applying these models, we carried out the first ultracool dwarf search with LSST DP1, cross-matched against VHS and CatWISE. Crucially, Euclid Quick Data Release 1 provided near-IR spectra for hundreds of candidates, enabling a rare, large-scale external spectroscopic validation. This confirmed 15 M6--L2 discoveries, verified USMILE performance, and clarified regimes where USMILE predictions are most reliable. Building on these insights, we identified 25 additional M6--L9 photometric candidates. These demonstrate the effectiveness of machine-learning methods in the data-rich era of wide-field surveys, highlighting the synergy between LSST and Euclid in expanding the ultracool dwarf census.

Ultracool dwarf Science with MachIne LEarning (USMILE). I. Scalable Tree-Based Models for Photometric Spectral Classification and New Discoveries from LSST Data Preview 1 and Euclid Quick Data Release 1

TL;DR

The paper tackles the challenge of discovering ultracool dwarfs in the data-rich LSST and Euclid era by building USMILE Avocado, a scalable two-model framework (classifier and regressor) based on gradient-boosted trees that can handle missing photometric data. It leverages a large, uncertainty-augmented labeled set derived from UltracoolSheet, reddened stars, and quasars, with eight colors anchored on computed from synthesized LSST, VHS, and CatWISE photometry, and validated against Euclid QDR1 spectra. The results show ROC AUC , F1 for classification, and regressor MSE subtypes, with external spectroscopic validation from Euclid confirming 15 new M6–L2 ultracool dwarfs and guiding reliability regimes, plus 25 additional candidates. The work demonstrates scalable, bias-resistant photometric spectral classification using ML in the LSST-Euclid era, enabling a more complete ultracool-dwarf census and informing follow-up strategies.

Abstract

We present the Ultracool dwarf Science with MachIne LEarning (USMILE), a program developing machine-learning tools for the discovery and characterization of ultracool dwarfs. We introduce USMILE Avocado, a spectral classification framework that uses broadband photometry from wide-field surveys -- Rubin Observatory LSST Data Preview 1, VISTA Hemisphere Survey, and CatWISE -- as input features. The framework has two gradient-boosted decision-tree models scalable to the massive data volumes of modern surveys: the classifier, which distinguishes ultracool dwarfs from stellar/extragalactic contaminants, and the regressor, which predicts spectral types. A key strength is its ability to natively handle missing photometric features, whereas earlier machine-learning approaches required complete multi-band detections or relied on imputation, thereby excluding genuine ultracool dwarfs or introducing bias. Trained on an augmented labeled dataset of >2 million sources built from known ultracool dwarfs, reddened early-type stars, and quasars, the models achieve strong performance: the classifier attains an ROC AUC of 0.976 and an F1 score of 0.92, while the regressor yields a mean-squared error of 0.88 subtypes. Applying these models, we carried out the first ultracool dwarf search with LSST DP1, cross-matched against VHS and CatWISE. Crucially, Euclid Quick Data Release 1 provided near-IR spectra for hundreds of candidates, enabling a rare, large-scale external spectroscopic validation. This confirmed 15 M6--L2 discoveries, verified USMILE performance, and clarified regimes where USMILE predictions are most reliable. Building on these insights, we identified 25 additional M6--L9 photometric candidates. These demonstrate the effectiveness of machine-learning methods in the data-rich era of wide-field surveys, highlighting the synergy between LSST and Euclid in expanding the ultracool dwarf census.
Paper Structure (8 sections, 3 figures)

This paper contains 8 sections, 3 figures.

Figures (3)

  • Figure 1: Photometric differences between Pan-STARRS1 and LSST in the $i$ (top), $z$ (middle), and $y$ (bottom) bands, derived from synthetic model spectra of the SPHINX (green), Sonora Diamondback (orange), and Exo-REM (blue) grids. Each circle marks an individual synthetic spectrum from the respective model grid.
  • Figure 2: The t-SNE projection of the complete labeled dataset (Section \ref{['subsec:classifier']}). Positive (ultracool dwarfs) and negative (reddened early-type stars and quasars) samples are shown in blue and grey, respectively.
  • Figure 3: ROC curves for the USMILE baseline classifier (left) and customized classifier (right). The corresponding area-under-the-curve (AUC) values are summarized in Table \ref{['tab:performance']}. Open circles mark the points corresponding to the selected probability threshold for binary classification. In the left panel, the false-positive rate on the x-axis is shown on a logarithmic scale.