Sequential monitoring for distributional changepoint using degenerate U-statistics
Cooper Boniece, Lajos Horvath, Lorenzo Trapani
TL;DR
The paper develops online changepoint detection for distributional changes in a sequence $\{\mathbf{X}_i\}$ using degenerate U-statistics, introducing weighted CUSUM and Page-CUSUM detectors and a novel repurposing scheme that incorporates past monitoring data into the training sample. It delivers a thorough asymptotic theory under $H_0$ and $H_A$, requiring only square-summability $\sum_{\ell} \lambda_{\ell}^2<\infty$ of kernel eigenvalues, and provides Monte Carlo methods to approximate critical values and to quantify detection delays for early and late changes. The work demonstrates strong empirical performance, particularly for multivariate data, and offers practical kernel choices (including energy distance and Grothendieck divergences) along with training-sample stability testing and moving-window extensions. Overall, it broadens online changepoint detection by relaxing kernel-spectral assumptions, enabling robust, distribution-wide monitoring with interpretable delay metrics and adaptable horizon strategies.
Abstract
We investigate the online detection of changepoints in the distribution of a sequence of observations using degenerate U-statistic-type processes. We study weighted versions of: an ordinary, CUSUM-type scheme, a Page-CUSUM-type scheme, and an entirely novel approach based on recycling past observations into the training sample. With an emphasis on completeness, we consider open-ended and closed-ended schemes, in the latter case considering both short- and long-running monitoring schemes. We study the asymptotics under the null in all cases, also proposing a consistent, Monte-Carlo based approximation of critical values; and we derive the limiting distribution of the detection delays under early and late occurring changes under the alternative, thus enabling to quantify the expected delay associated with each procedure. As a crucial technical contribution, we derive all our asymptotics under the assumption that the kernels associated with our U-statistics are square summable, instead of requiring the typical absolute summability, which makes our assumption naturally easier to check. Our simulations show that our procedures work well in all cases considered, having excellent power versus several types of distributional changes, and appearing to be particularly suited to the analysis of multivariate data.
