Beyond Distributional Fidelity: Causal-Penalized Diffusion for Synthetic Tabular Data
In this paper, we study whether causal fidelity can be improved directly within a fully generative tabular model.
Key points
- Synthetic tabular generators are commonly optimized for distributional fidelity, but statistical similarity alone does not guarantee preservation of causal effects.
- Causal Fidelity is defined with respect to a target estimand as the discrepancy between inferential distributions obtained from real and synthetic data, and theoretical results show that high statistical fidelity does not generally imply high causal fidelity.
- We then propose a causal-fidelity-aware training framework which adds a causal discrepancy penalty to the generative objective.
- We further establish conditions under which causal regularization improves expected causal fidelity.
Sources (1)
- [1]Beyond Distributional Fidelity: Causal-Penalized Diffusion for Synthetic Tabular DataarXiv (AI, ML, NLP, CV, robotics, multi-agent) · Oct 8, 07:36 AM
In this paper, we study whether causal fidelity can be improved directly within a fully generative tabular model.
Synthetic tabular generators are commonly optimized for distributional fidelity, but statistical similarity alone does not guarantee preservation of causal effects.
Extractive summary: sentences quoted from the sources.