Membership Inference over Diffusion-models-based Synthetic Tabular Data
Peini Cheng, Amir Bahmani
TL;DR
This work assesses privacy risks of diffusion-model-based synthetic tabular data by developing a step-wise error–based Membership Inference Attack (MIA) and applying it to two recent models, TabDDPM and TabSyn. It finds TabDDPM to be more vulnerable, particularly with smaller training sets, while TabSyn shows resilience against the proposed attacks. The authors critique standard privacy metrics like Distance to Closest Record (DCR) for diffusion models and advocate for targeted privacy evaluations to guide robust, privacy-preserving synthetic data design. The results underscore how model architecture and training data scale influence privacy risk, motivating future defenses and more comprehensive privacy testing frameworks in diffusion-based data synthesis.
Abstract
This study investigates the privacy risks associated with diffusion-based synthetic tabular data generation methods, focusing on their susceptibility to Membership Inference Attacks (MIAs). We examine two recent models, TabDDPM and TabSyn, by developing query-based MIAs based on the step-wise error comparison method. Our findings reveal that TabDDPM is more vulnerable to these attacks. TabSyn exhibits resilience against our attack models. Our work underscores the importance of evaluating the privacy implications of diffusion models and encourages further research into robust privacy-preserving mechanisms for synthetic data generation.
