Can we spot a fake?

Shahar Mendelson; Grigoris Paouris; Roman Vershynin

Can we spot a fake?

Shahar Mendelson, Grigoris Paouris, Roman Vershynin

Abstract

The problem of detecting fake data inspires the following seemingly simple mathematical question. Sample a data point $X$ from the standard normal distribution in $\mathbb{R}^n$. An adversary observes $X$ and corrupts it by adding a vector $rt$, where they can choose any vector $t$ from a fixed set $T$ of the adversary's "tricks", and where $r>0$ is a fixed radius. The adversary's choice of $t=t(X)$ may depend on the true data $X$. The adversary wants to hide the corruption by making the fake data $X+rt$ statistically indistinguishable from the real data $X$. What is the largest radius $r=r(T)$ for which the adversary can create an undetectable fake? We show that for highly symmetric sets $T$, the detectability radius $r(T)$ is approximately twice the scaled Gaussian width of $T$. The upper bound actually holds for arbitrary sets $T$ and generalizes to arbitrary, non-Gaussian distributions of real data $X$. The lower bound may fail for not highly symmetric $T$, but we conjecture that this problem can be solved by considering the focused version of the Gaussian width of $T$, which focuses on the most important directions of $T$.

Can we spot a fake?

Abstract

The problem of detecting fake data inspires the following seemingly simple mathematical question. Sample a data point

from the standard normal distribution in

. An adversary observes

and corrupts it by adding a vector

, where they can choose any vector

from a fixed set

of the adversary's "tricks", and where

is a fixed radius. The adversary's choice of

may depend on the true data

. The adversary wants to hide the corruption by making the fake data

statistically indistinguishable from the real data

. What is the largest radius

for which the adversary can create an undetectable fake? We show that for highly symmetric sets

, the detectability radius

is approximately twice the scaled Gaussian width of

. The upper bound actually holds for arbitrary sets

and generalizes to arbitrary, non-Gaussian distributions of real data

. The lower bound may fail for not highly symmetric

, but we conjecture that this problem can be solved by considering the focused version of the Gaussian width of

, which focuses on the most important directions of

Can we spot a fake?

Abstract

Can we spot a fake?

Abstract

Paper Structure

Table of Contents

Key Result

Theorems & Definitions (17)