Testing Imprecise Hypotheses
Lucas Kania, Tudor Manole, Larry Wasserman, Sivaraman Balakrishnan
TL;DR
This work develops a rigorous minimax framework for tolerant testing of imprecise hypotheses, where the null distribution is allowed a neighborhood of radius $ε_0$ and alternatives lie at distance at least $ε_1$. Focusing on Gaussian sequence, smooth Gaussian white-noise, and density models, it derives sharp upper and lower bounds on the critical separation and reveals regime structures—free tolerance, interpolation, and functional estimation—for various norms, notably $ℓ_1$ and smooth $ℓ_p$ norms. It shows the classical χ^2 statistic is suboptimal in tolerant settings and proposes practical tests, including plug-in and debiased plug-in statistics, that achieve minimax rates. The results connect to estimation theory via moment-matching dualities and extend to systematic-uncertainty contexts in high-energy physics, offering a robust framework for robust goodness-of-fit testing in scientific applications.
Abstract
Many scientific applications involve testing theories that are only partially specified. This task often amounts to testing the goodness-of-fit of a candidate distribution while allowing for reasonable deviations from it. The tolerant testing framework provides a systematic way of constructing such tests. Rather than testing the simple null hypothesis that data was drawn from a candidate distribution, a tolerant test assesses whether the data is consistent with any distribution that lies within a given neighborhood of the candidate. As this neighborhood grows, the tolerance to misspecification increases, while the power of the test decreases. In this work, we characterize the information-theoretic trade-off between the size of the neighborhood and the power of the test, in several canonical models. On the one hand, we characterize the optimal trade-off for tolerant testing in the Gaussian sequence model, under deviations measured in both smooth and non-smooth norms. On the other hand, we study nonparametric analogues of this problem in smooth regression and density models. Along the way, we establish the sub-optimality of the classical chi-squared statistic for tolerant testing, and study simple alternative hypothesis tests.
