Investigating the Effects of Point Source Injection Strategies on KMTNet Real/Bogus Classification
Dongjin Lee, Gregory S. H. Paek, Seo-Won Chang, Changwan Kim, Mankeun Jeong, Hongjae Moon, Seong-Heon Lee, Jae-Hun Jung, Myungshin Im
TL;DR
The paper addresses the challenge of training real/bogus RB classifiers for time-domain astronomy under data scarcity and class imbalance by comparing point-source injection strategies. It systematically evaluates Random Injection, Near Galaxy Injection, and a combined approach (including a distance-filtered variant) using KMTNet data and a simulation-to-reality framework, with testing on real GW follow-up observations. Results show RI is strong for asteroid/bogus discrimination but weak for galaxy-proximate transients, NGI improves transient recall near galaxies but increases false positives, and RI+NGI variants achieve better balance, with RI+NGI$^ ightarrow$dagger notably reducing false positives while preserving detection. The study highlights injection strategy as a critical factor for robust RB classifiers in GW follow-up campaigns and provides guidance for tailoring training data to environmental context and survey characteristics.
Abstract
Recently, machine learning-based real/bogus (RB) classifiers have demonstrated effectiveness in filtering out artifacts and identifying genuine transients in real-time astronomical surveys. However, the rarity of transient events and the extensive human labeling required for a large number of samples pose significant challenges in constructing training datasets for RB classification. Given these challenges, point source injection techniques, which inject simulated point sources into optical images, provide a promising solution. This paper presents the first detailed comparison of different point source injection strategies and their effects on classification performance within a simulation-to-reality framework. To this end, we first construct various training datasets based on Random Injection (RI), Near Galaxy Injection (NGI), and a combined approach by using the Korea Microlensing Telescope Network datasets. Subsequently, we train convolutional neural networks on simulated cutout samples and evaluate them on real, imbalanced datasets from gravitational wave follow-up observations for GW190814 and S230518h. Extensive experimental results show that RI excels at asteroid detection and bogus filtering but underperforms on transients occurring near galaxies (e.g., supernovae). In contrast, NGI is effective for detecting transients near galaxies but tends to misclassify variable stars as transients, resulting in a high false positive rate. The combined approach effectively handles these trade-offs, thereby balancing between detection rate and false positive rate. Our results emphasize the importance of point source injection strategy in developing robust RB classifiers for transient (or multi-messenger) follow-up campaigns.
