OpenLVLM-MIA: A Controlled Benchmark Revealing the Limits of Membership Inference Attacks on Large Vision-Language Models
Ryoto Miyamoto, Xin Fan, Fuyuko Kido, Tsuneo Matsumoto, Hayato Yamana
TL;DR
OpenLVLM-MIA exposes the critical problem that many MIA evaluations for LVLMs conflate dataset biases with true membership signals. By constructing a bias-controlled, ground-truth-annotated 6,000-image benchmark and evaluating a transparent OpenCLIP-LLaVA pipeline across three training stages, the study shows that state-of-the-art MIAs perform at chance under proper controls. The work highlights the need for distribution auditing, multimodal-specific attack strategies, and fully disclosed training data to enable reliable privacy assessments. It provides public datasets, tooling, and models to foster reproducibility and guide future privacy-preserving developments in vision-language systems.
Abstract
OpenLVLM-MIA is a new benchmark that highlights fundamental challenges in evaluating membership inference attacks (MIA) against large vision-language models (LVLMs). While prior work has reported high attack success rates, our analysis suggests that these results often arise from detecting distributional bias introduced during dataset construction rather than from identifying true membership status. To address this issue, we introduce a controlled benchmark of 6{,}000 images where the distributions of member and non-member samples are carefully balanced, and ground-truth membership labels are provided across three distinct training stages. Experiments using OpenLVLM-MIA demonstrated that the performance of state-of-the-art MIA methods approached chance-level. OpenLVLM-MIA, designed to be transparent and unbiased benchmark, clarifies certain limitations of MIA research on LVLMs and provides a solid foundation for developing stronger privacy-preserving techniques.
