Table of Contents
Fetching ...

Response to Discussions of "Causal and Counterfactual Views of Missing Data Models"

Razieh Nabi, Rohit Bhattacharya, Ilya Shpitser, James M. Robins

TL;DR

This paper addresses identifiability of complete-data distributions under MNAR by leveraging graphical models to connect causal inference and missing data. By reframing identifiability as identifying the joint distribution over counterfactuals $L^{(1)}$ and missingness $R$, and deriving a counterfactual $g$-formula, it shows nonparametric identification under Markov restrictions of an m-DAG whenever $p(R=1\mid L^{(1)})$ is identifiable from the observed data, with $p(l^{(1)}) = p(l,R=1)/p(R=1\mid l^{(1)})$. Key contributions include clarifying two notions of nonparametric identification, cataloging identification results for various m-DAGs, examining relationships to SWIGs and instrumental/ shadow-variable strategies, and addressing estimation and validation challenges. The work provides a conceptual bridge between causal and missing-data theory, avoids rank-preservation assumptions, and outlines a roadmap to integrate graphical identification, auxiliary information, censoring-based interventions, and robust sensitivity analysis for practical MNAR analysis.

Abstract

We are grateful to the discussants, Levis and Kennedy [2025], Luo and Geng [2025], Wang and van der Laan [2025], and Yang and Kim [2025], for their thoughtful comments on our paper (Nabi et al., 2025). In this rejoinder, we summarize our main contributions and respond to each discussion in turn.

Response to Discussions of "Causal and Counterfactual Views of Missing Data Models"

TL;DR

This paper addresses identifiability of complete-data distributions under MNAR by leveraging graphical models to connect causal inference and missing data. By reframing identifiability as identifying the joint distribution over counterfactuals and missingness , and deriving a counterfactual -formula, it shows nonparametric identification under Markov restrictions of an m-DAG whenever is identifiable from the observed data, with . Key contributions include clarifying two notions of nonparametric identification, cataloging identification results for various m-DAGs, examining relationships to SWIGs and instrumental/ shadow-variable strategies, and addressing estimation and validation challenges. The work provides a conceptual bridge between causal and missing-data theory, avoids rank-preservation assumptions, and outlines a roadmap to integrate graphical identification, auxiliary information, censoring-based interventions, and robust sensitivity analysis for practical MNAR analysis.

Abstract

We are grateful to the discussants, Levis and Kennedy [2025], Luo and Geng [2025], Wang and van der Laan [2025], and Yang and Kim [2025], for their thoughtful comments on our paper (Nabi et al., 2025). In this rejoinder, we summarize our main contributions and respond to each discussion in turn.
Paper Structure (7 sections, 1 equation, 4 figures)

This paper contains 7 sections, 1 equation, 4 figures.

Figures (4)

  • Figure 1: (a) Ignorable treatment model with SWIGs; (b) Its MCAR analogue, with a single relevant SWIG; (c) Conditionally ignorable treatment model with SWIGs; (d) Its MAR analogue, again with only one relevant SWIG.
  • Figure 2: (a) A possible SWIG representation of a self-censoring missing data model; (b) Stitching SWIGs into an observed-data model produces a cycle; (c) A SWIG that draws a distinction between the full data variable $L^{(1)}$, viewed as an unobserved confounder $U$ with extra restrictions, and the observed variable $L$ under an intervention where $R$ is set to $1$; (d) The full data graph equating the full data variable $L^{(1)}$ with an unobserved confounder $U$ with special structure.
  • Figure 3: (a) The block-parallel MNAR model, labeling missing variables as unmeasured for the purposes of constructing a SWIG; (b) The corresponding SWIG obtained from (a).
  • Figure 4: (a) Extension of Levis and Kennedy's example to have two missing variables. (b) Graph obtained after fixing $R_A$ and $R_Y$ in parallel. (c) SWIG obtained by splitting the treatment variable $A$ after already having fixed $R_A$ and $R_Y$.