Table of Contents
Fetching ...

Discovering Causal Relationships using Proxy Variables under Unmeasured Confounding

Yong Wu, Yanwei Fu, Shouyan Wang, Yizhou Wang, Xinwei Sun

TL;DR

This work tackles causal hypothesis testing under unmeasured confounding by leveraging proxy variables. It introduces a novel integral-equation identification framework that links $p(y|x)$ to $p(w|x)$ using a single negative control outcome, under completeness and mild regularity, and develops a kernel-based PMCR procedure with bootstrap-based inference. To address nonidentifiability in some regimes, it extends the approach to incorporate a second proxy (NCE) and derive a two-proxy testing method that restores identifiability. The method is validated through extensive simulations in continuous and discrete settings and applied to real data from ICU records and the World Values Survey, with results aligning with established findings. Overall, the paper provides a principled, nonparametric toolkit for testing causal relationships under unmeasured confounding that is applicable across data types and supported by rigorous asymptotic theory and practical implementations.

Abstract

Inferring causal relationships between variable pairs in the observational study is crucial but challenging, due to the presence of unmeasured confounding. While previous methods employed the negative controls to adjust for the confounding bias, they were either restricted to the discrete setting (i.e., all variables are discrete) or relied on strong assumptions for identification. To address these problems, we develop a general nonparametric approach that accommodates both discrete and continuous settings for testing causal hypothesis under unmeasured confounders. By using only a single negative control outcome (NCO), we establish a new identification result based on a newly proposed integral equation that links the outcome and NCO, requiring only the completeness and mild regularity conditions. We then propose a kernel-based testing procedure that is more efficient than existing moment-restriction methods. We derive the asymptotic level and power properties for our tests. Furthermore, we examine cases where our procedure using only NCO fails to achieve identification, and introduce a new procedure that incorporates a negative control exposure (NCE) to restore identifiability. We demonstrate the effectiveness of our approach through extensive simulations and real-world data from the Intensive Care Data and World Values Survey.

Discovering Causal Relationships using Proxy Variables under Unmeasured Confounding

TL;DR

This work tackles causal hypothesis testing under unmeasured confounding by leveraging proxy variables. It introduces a novel integral-equation identification framework that links to using a single negative control outcome, under completeness and mild regularity, and develops a kernel-based PMCR procedure with bootstrap-based inference. To address nonidentifiability in some regimes, it extends the approach to incorporate a second proxy (NCE) and derive a two-proxy testing method that restores identifiability. The method is validated through extensive simulations in continuous and discrete settings and applied to real data from ICU records and the World Values Survey, with results aligning with established findings. Overall, the paper provides a principled, nonparametric toolkit for testing causal relationships under unmeasured confounding that is applicable across data types and supported by rigorous asymptotic theory and practical implementations.

Abstract

Inferring causal relationships between variable pairs in the observational study is crucial but challenging, due to the presence of unmeasured confounding. While previous methods employed the negative controls to adjust for the confounding bias, they were either restricted to the discrete setting (i.e., all variables are discrete) or relied on strong assumptions for identification. To address these problems, we develop a general nonparametric approach that accommodates both discrete and continuous settings for testing causal hypothesis under unmeasured confounders. By using only a single negative control outcome (NCO), we establish a new identification result based on a newly proposed integral equation that links the outcome and NCO, requiring only the completeness and mild regularity conditions. We then propose a kernel-based testing procedure that is more efficient than existing moment-restriction methods. We derive the asymptotic level and power properties for our tests. Furthermore, we examine cases where our procedure using only NCO fails to achieve identification, and introduce a new procedure that incorporates a negative control exposure (NCE) to restore identifiability. We demonstrate the effectiveness of our approach through extensive simulations and real-world data from the Intensive Care Data and World Values Survey.
Paper Structure (59 sections, 53 theorems, 348 equations, 10 figures, 3 tables)

This paper contains 59 sections, 53 theorems, 348 equations, 10 figures, 3 tables.

Key Result

Proposition 1

Under condition assum.completeness and regularity conditions assum.regularity condition, there exists a $h(w,y) \in \mathcal{L}^2\{F(w)\}$ for all $y$, such that it solves the following integral equation for all $(y,u)$:

Figures (10)

  • Figure 1: Causal diagrams over $X,Y,U,W,Z$. $W$ (resp.$Z$) denote the negative control outcome (resp. exposure). The dotted line indicates its potential presence or absence.
  • Figure 2: The change of power across $\gamma_W$ in example \ref{['example:linear_gaussian_two_proxy']} (left) and in the nonlinear example (right).
  • Figure 3: Type-I error rate (left) and power rate (right) of our testing procedure and baseline methods in the single-proxy setting. The solid line reports the average value over 20 times, and the shaded area denotes the region $(\mathrm{mean}-\mathrm{std},\mathrm{mean}+\mathrm{std})$.
  • Figure 4: Type-I error rate (left) and power rate (right) of our procedure with PMCR and the first-moment method in example \ref{['example:linear_gaussian']}.
  • Figure 5: Type-I error rate (left) and power rate (right) of our procedure and the Miao's method in the discrete setting.
  • ...and 5 more figures

Theorems & Definitions (103)

  • Proposition 1
  • Remark 1
  • Theorem 1
  • Corollary 1
  • Corollary 2
  • Remark 2
  • Remark 3
  • Theorem 2
  • Theorem 2
  • Remark 4
  • ...and 93 more