Table of Contents
Fetching ...

On the identifiability of causal graphs with multiple environments

Francesco Montagna

TL;DR

The paper shows that causal graphs underlying nonlinear structural causal models become identifiable from data in two sufficiently different environments, provided the noise is Gaussian. It leverages a deep connection between SCMs and independent component analysis, using second-order derivatives of the log-likelihood to constrain the inverse Jacobian and recover the graph structure up to a permutation, which is then fixed by acyclicity. The contribution includes a novel identifiability proof, an algorithmic approach to recover the Jacobian support, and empirical validation on synthetic data demonstrating full graph recovery where pure observational data would fail. The results open avenues for causality-focused identifiability results in multi-environment settings and suggest relaxing Gaussianity and scaling the method to higher dimensions in future work.

Abstract

Causal discovery from i.i.d. observational data is known to be generally ill-posed. We demonstrate that if we have access to the distribution of a structural causal model, and additional data from only two environments that sufficiently differ in the noise statistics, the unique causal graph is identifiable. Notably, this is the first result in the literature that guarantees the entire causal graph recovery with a constant number of environments and arbitrary nonlinear mechanisms. Our only constraint is the Gaussianity of the noise terms; however, we propose potential ways to relax this requirement. Of interest on its own, we expand on the well-known duality between independent component analysis (ICA) and causal discovery; recent advancements have shown that nonlinear ICA can be solved from multiple environments, at least as many as the number of sources: we show that the same can be achieved for causal discovery while having access to much less auxiliary information.

On the identifiability of causal graphs with multiple environments

TL;DR

The paper shows that causal graphs underlying nonlinear structural causal models become identifiable from data in two sufficiently different environments, provided the noise is Gaussian. It leverages a deep connection between SCMs and independent component analysis, using second-order derivatives of the log-likelihood to constrain the inverse Jacobian and recover the graph structure up to a permutation, which is then fixed by acyclicity. The contribution includes a novel identifiability proof, an algorithmic approach to recover the Jacobian support, and empirical validation on synthetic data demonstrating full graph recovery where pure observational data would fail. The results open avenues for causality-focused identifiability results in multi-environment settings and suggest relaxing Gaussianity and scaling the method to higher dimensions in future work.

Abstract

Causal discovery from i.i.d. observational data is known to be generally ill-posed. We demonstrate that if we have access to the distribution of a structural causal model, and additional data from only two environments that sufficiently differ in the noise statistics, the unique causal graph is identifiable. Notably, this is the first result in the literature that guarantees the entire causal graph recovery with a constant number of environments and arbitrary nonlinear mechanisms. Our only constraint is the Gaussianity of the noise terms; however, we propose potential ways to relax this requirement. Of interest on its own, we expand on the well-known duality between independent component analysis (ICA) and causal discovery; recent advancements have shown that nonlinear ICA can be solved from multiple environments, at least as many as the number of sources: we show that the same can be achieved for causal discovery while having access to much less auxiliary information.
Paper Structure (43 sections, 12 theorems, 67 equations, 4 figures, 2 algorithms)

This paper contains 43 sections, 12 theorems, 67 equations, 4 figures, 2 algorithms.

Key Result

Proposition 1

Let $J_{\mathbf f^{-1}}(\mathbf x)$ faithful. Then, for each $i \neq j$:

Figures (4)

  • Figure 1: verage SHD ($0$ is best, $1$ is worst) achieved by \ref{['alg:jacobian_support_sketch']} over $50$ seeds on binary graphs. When the assumptions of \ref{['thm:identifiability']} are satisfied, the method can appropriately infer the causal direction, both in the observationally identifiable setting (nonlinear ANM, PNL, LSNM) and the observationally non-identifiable one (linear Gaussian model and the three SCMs with arbitrary nonlinearity). The number of environments does not have a notable effect on the accuracy.
  • Figure 2: We plot the Gamma density for different values of shape and scale. The left plot fixes the shape $\alpha=1$; the right plot fixes $\alpha=2$. We let $\theta$ vary to illustrate how the distribution changes between the rescaling environments of our experiments. We note that for $\alpha=1$ the density doesn't have a finite critical point.
  • Figure 3: Average SHD ($0$ is best, $1$ is worst) achieved by \ref{['alg:jacobian_support_sketch']} over $50$ seeds on binary graphs. The sources are sampled from a gamma distribution with $\alpha \in [0.5, 1]$. In line with our theory, when the sources are generated according to a density that doesn't have critical points, our algorithm generally fails to infer the causal direction.
  • Figure 4: Average SHD ($0$ is best, $1$ is worst) achieved by \ref{['alg:jacobian_support_sketch']} over $50$ seeds on binary graphs. The sources are sampled from a gamma distribution with $\alpha \in [2, 2.5]$, which guarantees at least one point where the gradient of the log-likelihood vanishes (see \ref{['fig:alpha2']}). Interestingly, this appears to enable accurate inference of the causal graph when the number of environments increases.

Theorems & Definitions (27)

  • Definition 1: ICA model
  • Definition 2: Faithfulness
  • Proposition 1: Proposition 1 in reizinger2023jacobianbased
  • Definition 3: Environment
  • Definition 4: Identifiability of the causal graph
  • Lemma 1
  • Theorem 1
  • proof : Proof sketch (full proof in \ref{['app:thm1_proof']})
  • Lemma 2: Full rank of $\Omega_l$ under rescalings
  • proof
  • ...and 17 more