Table of Contents
Fetching ...

The selection function of the Gaia DR3 open cluster census

Emily L. Hunt, Tristan Cantat-Gaudin, Friedrich Anders, Sagar Malhotra, Lorenzo Spina, Alfred Castro-Ginard, Lorenzo Cavallo

TL;DR

This work delivers the first global selection function for Gaia DR3 open clusters (valid for the HR24 catalogue) using a comprehensive injection–retrieval approach and a logistic model anchored on $n_*$, $\text{med}(\sigma_{\varpi})$, $\rho_{\text{data}}$, and a CST threshold. It pairs a detailed cluster simulator with an efficient Ga ia data-density map and compares ML and logistic formulations, achieving $\approx 94.5\%$ accuracy while using only a small set of interpretable inputs. The authors also provide an XGBoost predictor to rapidly estimate cluster observables ($n_*$ and $\text{med}(\sigma_{\varpi})$) and publish the full injection/retrieval data, memberships, and code in open-source form ($ocelot$), enabling selection-effect-corrected analyses and cross-comparisons with extragalactic cluster populations. This methodology lays the groundwork for robust, bias-aware studies of the Milky Way’s open-cluster census and its evolution, especially for future Gaia data releases.

Abstract

Open clusters are among the most useful and widespread tracers of Galactic structure. The completeness of the Galactic open cluster census, however, remains poorly understood. For the first time ever, we establish the selection function of an entire open cluster census, publishing our results as an open-source Python package for use by the community. Our work is valid for the Hunt & Reffert catalogue of clusters in Gaia DR3. We developed and open-sourced our cluster simulator from our first work. Then, we performed 80,590 injection and retrievals of simulated open clusters to test the Hunt & Reffert catalogue's sensitivity. We fit a logistic model of cluster detectability that depends only on a cluster's number of stars, median parallax error, Gaia data density, and a user-specified significance threshold. We find that our simple model accurately predicts cluster detectability, with a 94.53\% accuracy on our training data that is comparable to a machine-learning based model with orders of magnitude more parameters. Our model itself offers numerous insights on why certain clusters are detected. We briefly use our model to show that cluster detectability depends on non-intuitive parameters, such as a cluster's proper motion, and we show that even a modest 25 km/s boost to a cluster's orbital speed can result in an almost 3$\times$ higher detection probability, depending on its position. In addition, we publish our raw cluster injection and retrievals and cluster memberships, which could be used for a number of other science cases -- such as estimating cluster membership incompleteness. Using our results, selection effect-corrected studies are now possible with the open cluster census. Our work will enable a number of brand new types of study, such as detailed comparisons between the Milky Way's cluster census and recent extragalactic cluster samples.

The selection function of the Gaia DR3 open cluster census

TL;DR

This work delivers the first global selection function for Gaia DR3 open clusters (valid for the HR24 catalogue) using a comprehensive injection–retrieval approach and a logistic model anchored on , , , and a CST threshold. It pairs a detailed cluster simulator with an efficient Ga ia data-density map and compares ML and logistic formulations, achieving accuracy while using only a small set of interpretable inputs. The authors also provide an XGBoost predictor to rapidly estimate cluster observables ( and ) and publish the full injection/retrieval data, memberships, and code in open-source form (), enabling selection-effect-corrected analyses and cross-comparisons with extragalactic cluster populations. This methodology lays the groundwork for robust, bias-aware studies of the Milky Way’s open-cluster census and its evolution, especially for future Gaia data releases.

Abstract

Open clusters are among the most useful and widespread tracers of Galactic structure. The completeness of the Galactic open cluster census, however, remains poorly understood. For the first time ever, we establish the selection function of an entire open cluster census, publishing our results as an open-source Python package for use by the community. Our work is valid for the Hunt & Reffert catalogue of clusters in Gaia DR3. We developed and open-sourced our cluster simulator from our first work. Then, we performed 80,590 injection and retrievals of simulated open clusters to test the Hunt & Reffert catalogue's sensitivity. We fit a logistic model of cluster detectability that depends only on a cluster's number of stars, median parallax error, Gaia data density, and a user-specified significance threshold. We find that our simple model accurately predicts cluster detectability, with a 94.53\% accuracy on our training data that is comparable to a machine-learning based model with orders of magnitude more parameters. Our model itself offers numerous insights on why certain clusters are detected. We briefly use our model to show that cluster detectability depends on non-intuitive parameters, such as a cluster's proper motion, and we show that even a modest 25 km/s boost to a cluster's orbital speed can result in an almost 3 higher detection probability, depending on its position. In addition, we publish our raw cluster injection and retrievals and cluster memberships, which could be used for a number of other science cases -- such as estimating cluster membership incompleteness. Using our results, selection effect-corrected studies are now possible with the open cluster census. Our work will enable a number of brand new types of study, such as detailed comparisons between the Milky Way's cluster census and recent extragalactic cluster samples.
Paper Structure (19 sections, 6 equations, 15 figures, 5 tables)

This paper contains 19 sections, 6 equations, 15 figures, 5 tables.

Figures (15)

  • Figure 1: Example of a simulated cluster with photometry broadened by differential reddening. The cluster has an age of 1 Gyr, a mass of 500 M, an extinction $A_V$ of 1.5 mag, 0.2 mag of differential reddening (measured in colour $E(B-V)$), and is at a distance of 1 kpc. Top: random fractal noise map used to pick a differential reddening for each star. Member stars are shown as white circles, with the coloured background corresponding to changes in extinction $A_V$. Bottom: final simulated CMD for the cluster, including the impact of differential reddening, unresolved binary stars, and Gaia photometric errors. Member stars are colour-coded based on their differential reddening with the same colour scheme as the upper panel. The black line shows the isochrone that the cluster's member stars were originally drawn from. The labelled black arrow shows the mean extinction vector for stars in this cluster.
  • Figure 2: Top: binned mean of $f_\text{detected}$ against the number of simulated stars injected into Gaia data for simulated clusters in this study. See also Fig. \ref{['fig:f_detected_trends']} for more examples of the logistic dependence of $f_\text{detected}$ on $n_*$. Bottom: a 2D histogram of CST against number of injected stars for simulated clusters. The number of clusters at a given point is shown by the logarithmic colour scale. The dashed grey line shows $\text{CST}=3$, which is the threshold for a cluster to be detected and included in HR23. Only clusters with a CST below 38 are shown; values larger than 38 are calculated as infinity by our Python method due to floating point precision reasons.
  • Figure 3: Four maps of Gaia data density at various different proper motion and parallax levels. The proper motion and parallax values of each subplot are indicated in the titles, with proper motions given in mas yr$^{-1}$ and parallax in mas. Maps are plotted in a Mollweide projection in Galactic coordinates, with $l$ increasing to the left, north up, and the Galactic centre in the middle. Vertical and horizontal grid lines show spacings of 60 and 30 degrees in $l$ and $b$ respectively.
  • Figure 4: Mean absolute SHAP values of the XGBoost model used to investigate different selection function parameters in Sect. \ref{['sec:models:xgboost']}.
  • Figure 5: Comparison between binned $P(\text{detected})$ for various $n_*$ vs. $\tilde{R}_\rho$ values, along with predictions of our selection function model, given $\tau=3$ (the CST threshold for detection of HR23). Left: observed distribution of injected and retrieved clusters. Centre left: predictions from our model. Right: remaining residuals between the model and observed data, where blue corresponds to model overconfidence, red to model underconfidence, and grey to a correct model prediction.
  • ...and 10 more figures