Outlier-robust additive matrix decomposition

Philip Thompson

Outlier-robust additive matrix decomposition

Philip Thompson

TL;DR

This work introduces outlier-robust additive matrix decomposition (RTRMD) for least-squares trace regression when the parameter is the sum of a low-rank and a sparse matrix under adversarial label contamination. It develops three design properties—PP, IP, and MP—derived from product-process, Chevet, and multiplier inequalities to control the interplay between decomposition, contamination, and design-noise interaction. The proposed estimators use sorted-Huber losses with decomposable regularizers and achieve near-optimal rates $\mathsf{r}(n,d_{eff,r}) + \mathsf{r}(n,d_{eff,s}) + \sqrt{(1+\log(1/\delta))/n} + \varepsilon\log(1/\varepsilon)$, adaptively across $s,r,\varepsilon,\delta$, and with $\delta$-subgaussian guarantees even when noise depends on features. The paper also provides lower bounds and simulations confirming substantial improvements of sorted-Huber losses over classical Huber and validates practical performance in contaminated, high-dimensional settings.

Abstract

We study least-squares trace regression when the parameter is the sum of a $r$-low-rank matrix and a $s$-sparse matrix and a fraction $ε$ of the labels is corrupted. For subgaussian distributions and feature-dependent noise, we highlight three needed design properties, each one derived from a different process inequality: a "product process inequality", "Chevet's inequality" and a "multiplier process inequality". These properties handle, simultaneously, additive decomposition, label contamination and design-noise interaction. They imply the near-optimality of a tractable estimator with respect to the effective dimensions $d_{eff,r}$ and $d_{eff,s}$ of the low-rank and sparse components, $ε$ and the failure probability $δ$. The near-optimal rate is $\mathsf{r}(n,d_{eff,r}) + \mathsf{r}(n,d_{eff,s}) + \sqrt{(1+\log(1/δ))/n} + ε\log(1/ε)$, where $\mathsf{r}(n,d_{eff,r})+\mathsf{r}(n,d_{eff,s})$ is the optimal rate in average with no contamination. Our estimator is adaptive to $(s,r,ε,δ)$ and, for fixed absolute constant $c>0$, it attains the mentioned rate with probability $1-δ$ uniformly over all $δ\ge\exp(-cn)$. Without matrix decomposition, our analysis also entails optimal bounds for a robust estimator adapted to the noise variance. Our estimators are based on "sorted" versions of Huber's loss. We present simulations matching the theory. In particular, it reveals the superiority of "sorted" Huber's losses over the classical Huber's loss.

Outlier-robust additive matrix decomposition

TL;DR

, adaptively across

, and with

-subgaussian guarantees even when noise depends on features. The paper also provides lower bounds and simulations confirming substantial improvements of sorted-Huber losses over classical Huber and validates practical performance in contaminated, high-dimensional settings.

Abstract

We study least-squares trace regression when the parameter is the sum of a

-low-rank matrix and a

-sparse matrix and a fraction

of the labels is corrupted. For subgaussian distributions and feature-dependent noise, we highlight three needed design properties, each one derived from a different process inequality: a "product process inequality", "Chevet's inequality" and a "multiplier process inequality". These properties handle, simultaneously, additive decomposition, label contamination and design-noise interaction. They imply the near-optimality of a tractable estimator with respect to the effective dimensions

and

of the low-rank and sparse components,

and the failure probability

. The near-optimal rate is

, where

is the optimal rate in average with no contamination. Our estimator is adaptive to

and, for fixed absolute constant

, it attains the mentioned rate with probability

uniformly over all

. Without matrix decomposition, our analysis also entails optimal bounds for a robust estimator adapted to the noise variance. Our estimators are based on "sorted" versions of Huber's loss. We present simulations matching the theory. In particular, it reveals the superiority of "sorted" Huber's losses over the classical Huber's loss.

Paper Structure (45 sections, 45 theorems, 296 equations, 5 figures)

This paper contains 45 sections, 45 theorems, 296 equations, 5 figures.

Introduction
General framework
Our estimators
Motivating examples
Results for RTRMD and robust sparse/low-rank regression
Related work and contributions
Robust sparse regression
Robust trace regression
Robust trace regression with additive matrix decomposition
The multiplier and product processes
Properties for subgaussian distributions
$\mathop{\mathrm{MP}}\limits$ in $M$-estimation with decomposable regularizers
First general theorem
Second general theorem
Simulation results
...and 30 more sections

Key Result

Theorem 2

Grant Assumptions assump:label:contamination-assump:distribution:subgaussian, model equation:structural:equation and assume $\mathbf{X}$ is isotropic. Then there are absolute constants $\sc\in(0,1)$, $c_1\in(0,1/2)$ and $C_0,C_1\ge1$ such that the following holds. Suppose $C_1^2L^4\epsilon\log(1/\ep Then, for any $\delta\in(0,1)$ such that $\delta\ge \exp\left(-\frac{n}{C_0L^4}\right),$ on an even

Figures (5)

Figure 1: Huber vs "Sorted" Huber losses in sparse regression: $\sqrt{\texttt{MSE}}$ versus $\epsilon$.
Figure 2: Robust sparse linear regression: different levels of sparsity (a) and comparisons between methods (b).
Figure 3: Robust low-rank trace regression: different values of rank (a), and comparisons between methods (b,c).
Figure 4: Trace regression with additive matrix decomposition with $\epsilon=0$
Figure 5: Robust trace regression with additive decomposition: different values of rank/sparsity (a), and comparisons between methods (b,c).

Theorems & Definitions (83)

Definition 1: Sorted Huber-type losses
Remark 1: Huber loss versus "Sorted" Huber loss
Theorem 2: Robust trace regression with additive matrix decomposition
Proposition 1
Theorem 3: Robust sparse/low-rank regression
Theorem 4: $\sigma$-adaptive robust sparse/low-rank regression
Remark 2
Remark 3: Comparison with 2019chinot
Remark 4: Comparison with 2019dalalyan:thompson
Remark 5: Further references in robust sparse regression
...and 73 more

Outlier-robust additive matrix decomposition

TL;DR

Abstract

Outlier-robust additive matrix decomposition

Authors

TL;DR

Abstract

Table of Contents

Key Result

Figures (5)

Theorems & Definitions (83)