ThreatIntel-Andro: Expert-Verified Benchmarking for Robust Android Malware Research
Hongpeng Bai, Minhong Dong, Yao Zhang, Shunzhe Zhao, Haobo Zhang, Lingyue Li, Yude Bai, Guangquan Xu
TL;DR
ThreatIntel-Andro addresses label noise and temporal drift in Android malware benchmarks by creating a fully expert-verified, traceable dataset linking samples to professional analysis reports from major vendors. It collects 5,123 samples across 146 families with IoCs and temporal anchors (2016–2025), linked to Koodous, and augments with static and dynamic features to enable robust behavioral analysis. Experiments show classifiers trained on ThreatIntel-Andro outperform those trained on engine-based labels, while simulated and real-world label noise degrade performance, especially for meta-learning approaches. The dataset provides a reliable baseline for open-set recognition and concept drift studies, enabling reproducible, credible malware research and adversarial robustness analyses.
Abstract
The rapidly evolving Android malware ecosystem demands high-quality, real-time datasets as a foundation for effective detection and defense. With the widespread adoption of mobile devices across industrial systems, they have become a critical yet often overlooked attack surface in industrial cybersecurity. However, mainstream datasets widely used in academia and industry (e.g., Drebin) exhibit significant limitations: on one hand, their heavy reliance on VirusTotal's multi-engine aggregation results introduces substantial label noise; on the other hand, outdated samples reduce their temporal relevance. Moreover, automated labeling tools (e.g., AVClass2) suffer from suboptimal aggregation strategies, further compounding labeling errors and propagating inaccuracies throughout the research community.
