Blind Baselines Beat Membership Inference Attacks for Foundation Models

Debeshee Das; Jie Zhang; Florian Tramèr

Blind Baselines Beat Membership Inference Attacks for Foundation Models

Debeshee Das, Jie Zhang, Florian Tramèr

TL;DR

This work shows that MI evaluations for foundation models are flawed because member/non-member data are often drawn from different distributions. It introduces simple blind attacks—date Detection, bag-of-words, and greedy n-gram methods—that ignore the model yet outperform published MI attacks across eight datasets, highlighting pervasive distribution shifts including temporal, replication, and tail differences. The study systematically analyzes case studies (WikiMIA, BookMIA, Temporal Wiki/arXiv, ArXiv variants, LAION-MI, and Gutenberg) and demonstrates that current evaluations can even be worse than random chance. It then argues for a shift toward IID and well-curated benchmarks (e.g., Pile, DataComp/DataComp-LM) and provides reproducible code to enable credible MI assessment in future work.

Abstract

Membership inference (MI) attacks try to determine if a data sample was used to train a machine learning model. For foundation models trained on unknown Web data, MI attacks are often used to detect copyrighted training materials, measure test set contamination, or audit machine unlearning. Unfortunately, we find that evaluations of MI attacks for foundation models are flawed, because they sample members and non-members from different distributions. For 8 published MI evaluation datasets, we show that blind attacks -- that distinguish the member and non-member distributions without looking at any trained model -- outperform state-of-the-art MI attacks. Existing evaluations thus tell us nothing about membership leakage of a foundation model's training data.

Blind Baselines Beat Membership Inference Attacks for Foundation Models

TL;DR

Abstract

Paper Structure (44 sections, 1 figure, 4 tables)

This paper contains 44 sections, 1 figure, 4 tables.

Introduction
Background and Related Work
Web-scale training datasets.
Membership inference attacks.
Membership inference for foundation models.
Evaluating membership inference.
Blindly Inferring Membership
Distribution Shifts in MI Evaluation Datasets
Temporal shifts.
Biases in data replication.
Distinguishable tails.
Blind Attack Techniques
Date detection.
Bag-of-words classification.
Greedy rare word selection.
...and 29 more sections

Figures (1)

Figure 1: In some cases, the distribution shift is easy to visualise. Here is the PCA plot of the WikiMIA dataset which shows members and non-members forming distinguishable clusters.

Blind Baselines Beat Membership Inference Attacks for Foundation Models

TL;DR

Abstract

Blind Baselines Beat Membership Inference Attacks for Foundation Models

Authors

TL;DR

Abstract

Table of Contents

Figures (1)