Replicable Clustering

Hossein Esfandiari; Amin Karbasi; Vahab Mirrokni; Grigoris Velegkas; Felix Zhou

Replicable Clustering

Hossein Esfandiari, Amin Karbasi, Vahab Mirrokni, Grigoris Velegkas, Felix Zhou

TL;DR

The paper studies replicable clustering under the distributional learning framework, introducing replicable algorithms for statistical $k$-means, $k$-medians, and $k$-centers. It builds replicable coresets via a replicable quad-tree, ensuring the cost on the coreset approximates the population cost, and leverages black-box approximation oracles to obtain near-optimal partitions with high probability. For Euclidean data, it further employs Johnson-Lindenstrauss dimensionality reduction to achieve poly$(d)$ sample complexity and outputs a clustering function that labels points rather than explicitly outputting centers. The results establish a principled path to replicable clustering with strong utility guarantees, supported by experiments on synthetic 2D distributions. This advances reliable, repeatable unsupervised learning in high dimensions by combining coresets, stable discretizations, and dimensionality reduction with robust probabilistic guarantees.

Abstract

We design replicable algorithms in the context of statistical clustering under the recently introduced notion of replicability from Impagliazzo et al. [2022]. According to this definition, a clustering algorithm is replicable if, with high probability, its output induces the exact same partition of the sample space after two executions on different inputs drawn from the same distribution, when its internal randomness is shared across the executions. We propose such algorithms for the statistical $k$-medians, statistical $k$-means, and statistical $k$-centers problems by utilizing approximation routines for their combinatorial counterparts in a black-box manner. In particular, we demonstrate a replicable $O(1)$-approximation algorithm for statistical Euclidean $k$-medians ($k$-means) with $\operatorname{poly}(d)$ sample complexity. We also describe an $O(1)$-approximation algorithm with an additional $O(1)$-additive error for statistical Euclidean $k$-centers, albeit with $\exp(d)$ sample complexity. In addition, we provide experiments on synthetic distributions in 2D using the $k$-means++ implementation from sklearn as a black-box that validate our theoretical results.

Replicable Clustering

TL;DR

The paper studies replicable clustering under the distributional learning framework, introducing replicable algorithms for statistical

-means,

-medians, and

-centers. It builds replicable coresets via a replicable quad-tree, ensuring the cost on the coreset approximates the population cost, and leverages black-box approximation oracles to obtain near-optimal partitions with high probability. For Euclidean data, it further employs Johnson-Lindenstrauss dimensionality reduction to achieve poly

sample complexity and outputs a clustering function that labels points rather than explicitly outputting centers. The results establish a principled path to replicable clustering with strong utility guarantees, supported by experiments on synthetic 2D distributions. This advances reliable, repeatable unsupervised learning in high dimensions by combining coresets, stable discretizations, and dimensionality reduction with robust probabilistic guarantees.

Abstract

-medians, statistical

-means, and statistical

-centers problems by utilizing approximation routines for their combinatorial counterparts in a black-box manner. In particular, we demonstrate a replicable

-approximation algorithm for statistical Euclidean

-medians (

-means) with

sample complexity. We also describe an

-approximation algorithm with an additional

-additive error for statistical Euclidean

-centers, albeit with

sample complexity. In addition, we provide experiments on synthetic distributions in 2D using the

-means++ implementation from sklearn as a black-box that validate our theoretical results.

Paper Structure (47 sections, 55 theorems, 193 equations, 2 figures, 1 table, 7 algorithms)

This paper contains 47 sections, 55 theorems, 193 equations, 2 figures, 1 table, 7 algorithms.

Introduction
Related Works
Setting & Notation
Clustering Methods and Generalizations
Parameters p and kappa
Replicability
Main Results
Overview of (k, p)-Clustering
Coresets
Replicable Quad Tree
Putting it Together
The Euclidean Metric, Dimensionality Reduction, and (k, p)-Clustering
Running Time for (k, p)-Clustering
Replicable k-Centers
Experiments
...and 32 more sections

Key Result

Theorem 3.1

Let $\varepsilon, \rho\in (0, 1)$. Given black-box access to a $\beta$-approximation oracle for weighted $k$-medians, respectively weighted $k$-means (cf. prob:cluster), there is a $\rho$-replicable algorithm for statistical $k$-medians, respectively $k$-means (cf. prob:stat cluster), such that with

Figures (2)

Figure 8.1: The results of running vanilla vs replicable $k$-Means++ on the two moons distribution for $k=3$.
Figure 8.2: The results of running vanilla vs replicable $k$-Means++ on a mixture of truncated Gaussians distributions for $k=3$.

Theorems & Definitions (61)

Definition 2.5: Replicable Algorithm; impagliazzo2022reproducibility
Theorem 3.1: Informal
Theorem 3.2: Informal
Theorem 3.3: Informal
Definition 4.1: (Strong) Coresets
Theorem 4.2: \ref{['thm:statistical clustering informal']}; Formal
Theorem 4.3: \ref{['thm:statistical clustering informal']}; Formal
Theorem 5.1: \ref{['thm:euclidean clustering informal']}; Formal
Proposition B.1: Bretagnolle-Huber-Carol Inequality; vaart1997weak
Remark B.2: Weighted $k$-Means/Medians
...and 51 more

Replicable Clustering

TL;DR

Abstract

Replicable Clustering

Authors

TL;DR

Abstract

Table of Contents

Key Result

Figures (2)

Theorems & Definitions (61)