Predict Training Data Quality via Its Geometry in Metric Space
Yang Ba, Mohammad Sadeq Abolhasani, Rong Pan
TL;DR
The paper addresses how training data geometry in a metric space affects AI model generalization. It introduces PH-based diversity measures derived from lifetimes in Vietoris–Rips filtrations, including $p_i = l_i/L$, Rényi persistence entropy $PE_k^{(q)}$, and PH-based Hill numbers $ ext{PEH}_k^{q}$, focusing on $H_0$ and $H_1$. It demonstrates that PH-based diversity satisfies key axioms and correlates with higher accuracy and lower variance in transfer-learning tasks with BERT across several text datasets, while entropy-based metrics like the Vendi Score fail to predict data quality. The results offer practical guidance for data selection and augmentation, showing that well-structured geometric diversity enables near-full-data performance with substantially smaller training sets.
Abstract
High-quality training data is the foundation of machine learning and artificial intelligence, shaping how models learn and perform. Although much is known about what types of data are effective for training, the impact of the data's geometric structure on model performance remains largely underexplored. We propose that both the richness of representation and the elimination of redundancy within training data critically influence learning outcomes. To investigate this, we employ persistent homology to extract topological features from data within a metric space, thereby offering a principled way to quantify diversity beyond entropy-based measures. Our findings highlight persistent homology as a powerful tool for analyzing and enhancing the training data that drives AI systems.
