How Realistic Is Your Synthetic Data? Constraining Deep Generative Models for Tabular Data

Mihaela Cătălina Stoian; Salijona Dyrmishi; Maxime Cordy; Thomas Lukasiewicz; Eleonora Giunchiglia

How Realistic Is Your Synthetic Data? Constraining Deep Generative Models for Tabular Data

Mihaela Cătălina Stoian, Salijona Dyrmishi, Maxime Cordy, Thomas Lukasiewicz, Eleonora Giunchiglia

TL;DR

This work tackles the problem that deep generative models for tabular data often violate domain knowledge encoded as linear inequalities. It introduces Constrained Deep Generative Models (C-DGMs) by automatically parsing linear constraints into a differentiable Constraint Layer (CL) that guarantees compliance and can minimally alter unconstrained outputs. Empirical results show standard DGMs exhibit high constraint violations, while C-DGMs substantially improve data utility and detection metrics across multiple datasets and models, with negligible impact on generation time. Additionally, CL can serve as an inference-time guardrail, offering flexibility when retraining is not feasible, and the approach advances safer, more reliable synthetic data generation for downstream ML tasks.

Abstract

Deep Generative Models (DGMs) have been shown to be powerful tools for generating tabular data, as they have been increasingly able to capture the complex distributions that characterize them. However, to generate realistic synthetic data, it is often not enough to have a good approximation of their distribution, as it also requires compliance with constraints that encode essential background knowledge on the problem at hand. In this paper, we address this limitation and show how DGMs for tabular data can be transformed into Constrained Deep Generative Models (C-DGMs), whose generated samples are guaranteed to be compliant with the given constraints. This is achieved by automatically parsing the constraints and transforming them into a Constraint Layer (CL) seamlessly integrated with the DGM. Our extensive experimental analysis with various DGMs and tasks reveals that standard DGMs often violate constraints, some exceeding $95\%$ non-compliance, while their corresponding C-DGMs are never non-compliant. Then, we quantitatively demonstrate that, at training time, C-DGMs are able to exploit the background knowledge expressed by the constraints to outperform their standard counterparts with up to $6.5\%$ improvement in utility and detection. Further, we show how our CL does not necessarily need to be integrated at training time, as it can be also used as a guardrail at inference time, still producing some improvements in the overall performance of the models. Finally, we show that our CL does not hinder the sample generation time of the models.

How Realistic Is Your Synthetic Data? Constraining Deep Generative Models for Tabular Data

TL;DR

Abstract

non-compliance, while their corresponding C-DGMs are never non-compliant. Then, we quantitatively demonstrate that, at training time, C-DGMs are able to exploit the background knowledge expressed by the constraints to outperform their standard counterparts with up to

improvement in utility and detection. Further, we show how our CL does not necessarily need to be integrated at training time, as it can be also used as a guardrail at inference time, still producing some improvements in the overall performance of the models. Finally, we show that our CL does not hinder the sample generation time of the models.

Paper Structure (34 sections, 7 theorems, 7 equations, 6 figures, 31 tables)

This paper contains 34 sections, 7 theorems, 7 equations, 6 figures, 31 tables.

Introduction
Problem Statement
Constrained Deep Generative Models
Experimental Analysis
Experimental Analysis Settings
Background Knowledge Alignment
Synthetic Data Quality
Post-processing Ability
Impact on Samples Generation Time
Related Work
Discussion and Conclusions
Theorems
Proof of Theorem \ref{['th:satisfaction']}
Proof of Theorem \ref{['th:refl']}
Proof of Theorem \ref{['th:min_change']}
...and 19 more sections

Key Result

Theorem 3.3

Let $m$ be a deep generative model. Let $\Pi$ be a satisfiable and finite set of constraints over the sample space. Let $\lambda$ be a variable ordering and $\text{CL}$ be the constraint layer built from $\lambda$ and $\Pi$. Then, the model C-$m$ obtained by incorporating $\text{CL}$ in $m$ is compl

Figures (6)

Figure 1: Overview on how to integrate $\text{CL}$ into a GAN-based model.
Figure 2: $\text{CL}$ in Example \ref{['ex:value_change']}.
Figure 3: Real data and samples generated by TableGAN and C-TableGAN for WiDS.
Figure 4: Samples generated by WGAN, CTGAN and TVAE, GOGGLE and their constrained versions for WiDS. The real and TableGAN distributions are given in Section \ref{['subsec:exp_violations']} in the main text.
Figure 5: Real data and samples generated by WGAN, TableGAN, CTGAN, TVAE, GOGGLE and their constrained versions for Heloc.
...and 1 more figures

Theorems & Definitions (14)

Example 2.1
Example 3.1
Example 3.2
Theorem 3.3
Corollary 3.4
Theorem 3.5
Theorem 3.6
proof
Lemma A.1
proof
...and 4 more

How Realistic Is Your Synthetic Data? Constraining Deep Generative Models for Tabular Data

TL;DR

Abstract

How Realistic Is Your Synthetic Data? Constraining Deep Generative Models for Tabular Data

Authors

TL;DR

Abstract

Table of Contents

Key Result

Figures (6)

Theorems & Definitions (14)