Table of Contents
Fetching ...

Incomplete U-Statistics of Equireplicate Designs: Berry-Esseen Bound and Efficient Construction

Cesare Miglioli, Jordan Awan

Abstract

U-statistics are a fundamental class of estimators that generalize the sample mean and underpin much of nonparametric statistics. Although extensively studied in both statistics and probability, key challenges remain: their high computational cost - addressed partly through incomplete U-statistics - and their non-standard asymptotic behavior in the degenerate case, which typically requires resampling methods for hypothesis testing. This paper presents a novel perspective on U-statistics, grounded in hypergraph theory and combinatorial designs. Our approach bypasses the traditional Hoeffding decomposition, the main analytical tool in this literature but one highly sensitive to degeneracy. By characterizing the dependence structure of a U-statistic, we derive a Berry-Esseen bound valid for incomplete U-statistics of deterministic designs, yielding conditions under which Gaussian limiting distributions can be established even in degenerate cases and when the order diverges. We also introduce efficient algorithms to construct incomplete U-statistics of equireplicate designs, a subclass of deterministic designs that, in certain cases, achieve minimum variance. Finally, we apply our framework to kernel-based tests that use Maximum Mean Discrepancy (MMD) and Hilbert-Schmidt Independence Criterion. In a real data example with the CIFAR-10 dataset, our permutation-free MMD test delivers substantial computational gains while retaining power and type I error control.

Incomplete U-Statistics of Equireplicate Designs: Berry-Esseen Bound and Efficient Construction

Abstract

U-statistics are a fundamental class of estimators that generalize the sample mean and underpin much of nonparametric statistics. Although extensively studied in both statistics and probability, key challenges remain: their high computational cost - addressed partly through incomplete U-statistics - and their non-standard asymptotic behavior in the degenerate case, which typically requires resampling methods for hypothesis testing. This paper presents a novel perspective on U-statistics, grounded in hypergraph theory and combinatorial designs. Our approach bypasses the traditional Hoeffding decomposition, the main analytical tool in this literature but one highly sensitive to degeneracy. By characterizing the dependence structure of a U-statistic, we derive a Berry-Esseen bound valid for incomplete U-statistics of deterministic designs, yielding conditions under which Gaussian limiting distributions can be established even in degenerate cases and when the order diverges. We also introduce efficient algorithms to construct incomplete U-statistics of equireplicate designs, a subclass of deterministic designs that, in certain cases, achieve minimum variance. Finally, we apply our framework to kernel-based tests that use Maximum Mean Discrepancy (MMD) and Hilbert-Schmidt Independence Criterion. In a real data example with the CIFAR-10 dataset, our permutation-free MMD test delivers substantial computational gains while retaining power and type I error control.
Paper Structure (44 sections, 18 theorems, 48 equations, 6 figures, 1 table, 3 algorithms)

This paper contains 44 sections, 18 theorems, 48 equations, 6 figures, 1 table, 3 algorithms.

Key Result

Proposition 1

$L(\mathcal{D})$ is the dependency graph of $\left\{ h(S), \;S \in D \right\}$.

Figures (6)

  • Figure 1: (Left) KS distance between the empirical distribution of the standardized incomplete uMMD statistics under $H_0$ and the $\mathcal{N}(0,1)$ distribution. (Right) Variance ratio under $H_1$ between the incomplete uMMD statistic based on equireplicate designs and that based on random designs. $95\%$ Monte Carlo CIs are included for both experiments.
  • Figure 2: (Left) KS distance between the empirical distribution of the standardized incomplete uHSIC statistics under $H_0$ and the $\mathcal{N}(0,1)$ distribution. (Right) Variance ratio under $H_1$ between the incomplete uHSIC statistic based on equireplicate designs and that based on random designs. $95\%$ Monte Carlo CIs are included for both experiments.
  • Figure 3: $95\%$ Monte Carlo CI for the power (left) and type I error (right) of the permutation-free (PF) version of the MMD test compared with its permutation-based (PB) counterpart, both evaluated on CIFAR-10 for different values of $n$ and $r$.
  • Figure 4: $1$-Equireplicate Partition for $n=6$ and its representation as a proper edge coloring of $K^{(2)}_6$. The construction of the matchings can be understood by holding "6" fixed, and rotating the other numbers clock-wise, which is equivalent to the construction in Theorem \ref{['thm:evenPartition']}.
  • Figure 5: $2$-Equireplicate Partition for $n=7$ and its representation as a decomposition of $K^{(2)}_7$ into disjoint cycles. The construction of the partition is done by holding the left column fixed and "rotating" the right column vertically, which is equivalent to the construction in Theorem \ref{['thm:oddPartition']}.
  • ...and 1 more figures

Theorems & Definitions (48)

  • Definition 1
  • Proposition 1
  • Lemma 1
  • Theorem 1: Berry-Esseen for Deterministic Designs
  • Corollary 1
  • Remark 1
  • Remark 2
  • Theorem 2
  • Example 1: Designs with unbalanced degree distribution
  • Corollary 2: CLT degenerate case of order $k-1$ for equireplicate designs
  • ...and 38 more