On the identifiability of causal graphs with multiple environments
Francesco Montagna
TL;DR
The paper shows that causal graphs underlying nonlinear structural causal models become identifiable from data in two sufficiently different environments, provided the noise is Gaussian. It leverages a deep connection between SCMs and independent component analysis, using second-order derivatives of the log-likelihood to constrain the inverse Jacobian and recover the graph structure up to a permutation, which is then fixed by acyclicity. The contribution includes a novel identifiability proof, an algorithmic approach to recover the Jacobian support, and empirical validation on synthetic data demonstrating full graph recovery where pure observational data would fail. The results open avenues for causality-focused identifiability results in multi-environment settings and suggest relaxing Gaussianity and scaling the method to higher dimensions in future work.
Abstract
Causal discovery from i.i.d. observational data is known to be generally ill-posed. We demonstrate that if we have access to the distribution of a structural causal model, and additional data from only two environments that sufficiently differ in the noise statistics, the unique causal graph is identifiable. Notably, this is the first result in the literature that guarantees the entire causal graph recovery with a constant number of environments and arbitrary nonlinear mechanisms. Our only constraint is the Gaussianity of the noise terms; however, we propose potential ways to relax this requirement. Of interest on its own, we expand on the well-known duality between independent component analysis (ICA) and causal discovery; recent advancements have shown that nonlinear ICA can be solved from multiple environments, at least as many as the number of sources: we show that the same can be achieved for causal discovery while having access to much less auxiliary information.
