Table of Contents
Fetching ...

Estimation of causal dose-response functions under data fusion

Jaewon Lim, Alex Luedtke

TL;DR

This work addresses estimating causal dose-response functions when data come from multiple partially aligned sources by formulating a data-fusion framework and deriving a Neyman-orthogonal loss for robust estimation. It presents two estimation strategies, including a kernel ridge regression method with a closed-form solution, and proves oracle excess-risk bounds that tighten with additional sources. The theory demonstrates that data fusion can reduce Lipschitz constants and improve worst-case performance under suitable eigenvalue decay, with minimax lower bounds supporting the advantage. Empirical results across diverse CDRF shapes and reference measures show consistent accuracy gains from data fusion, underscoring its practical value for non-smooth, function-valued causal parameters. The framework thus broadens the applicability of data fusion to causal inference tasks beyond standard averages, offering scalable and robust estimation in multi-source settings.

Abstract

Estimating the causal dose-response function is challenging, particularly when data from a single source are insufficient to estimate responses precisely across all exposure levels. To overcome this limitation, we propose a data fusion framework that leverages multiple data sources that are partially aligned with the target distribution. Specifically, we derive a Neyman-orthogonal loss function tailored for estimating the dose-response function within data fusion settings. To improve computational efficiency, we propose a stochastic approximation that retains orthogonality. We apply kernel ridge regression with this approximation, which provides closed-form estimators. Our theoretical analysis demonstrates that incorporating additional data sources yields tighter finite-sample regret bounds and improved worst-case performance, as confirmed via minimax lower bound comparison. Simulation studies validate the practical advantages of our approach, showing improved estimation accuracy when employing data fusion. This study highlights the potential of data fusion for estimating non-smooth parameters such as causal dose-response functions.

Estimation of causal dose-response functions under data fusion

TL;DR

This work addresses estimating causal dose-response functions when data come from multiple partially aligned sources by formulating a data-fusion framework and deriving a Neyman-orthogonal loss for robust estimation. It presents two estimation strategies, including a kernel ridge regression method with a closed-form solution, and proves oracle excess-risk bounds that tighten with additional sources. The theory demonstrates that data fusion can reduce Lipschitz constants and improve worst-case performance under suitable eigenvalue decay, with minimax lower bounds supporting the advantage. Empirical results across diverse CDRF shapes and reference measures show consistent accuracy gains from data fusion, underscoring its practical value for non-smooth, function-valued causal parameters. The framework thus broadens the applicability of data fusion to causal inference tasks beyond standard averages, offering scalable and robust estimation in multi-source settings.

Abstract

Estimating the causal dose-response function is challenging, particularly when data from a single source are insufficient to estimate responses precisely across all exposure levels. To overcome this limitation, we propose a data fusion framework that leverages multiple data sources that are partially aligned with the target distribution. Specifically, we derive a Neyman-orthogonal loss function tailored for estimating the dose-response function within data fusion settings. To improve computational efficiency, we propose a stochastic approximation that retains orthogonality. We apply kernel ridge regression with this approximation, which provides closed-form estimators. Our theoretical analysis demonstrates that incorporating additional data sources yields tighter finite-sample regret bounds and improved worst-case performance, as confirmed via minimax lower bound comparison. Simulation studies validate the practical advantages of our approach, showing improved estimation accuracy when employing data fusion. This study highlights the potential of data fusion for estimating non-smooth parameters such as causal dose-response functions.
Paper Structure (32 sections, 10 theorems, 176 equations, 1 figure, 1 table, 3 algorithms)

This paper contains 32 sections, 10 theorems, 176 equations, 1 figure, 1 table, 3 algorithms.

Key Result

Lemma 1

Assume cond:identifiabilitycond:target_modelcond:nuisances hold. If $\widehat{\theta}$ and $\widehat{g}$ are the target estimator and nuisance estimators from alg:gen_est, then

Figures (1)

  • Figure 1: Number line demonstrating approach for establishing data fusion necessarily improves worst-case performance. The key idea is to show that a high-probability worst-case upper bound (UB) of the risk for kernel ridge regression with data fusion lies below a high-probability worst-case lower bound (LB) for any estimator without it.

Theorems & Definitions (20)

  • Lemma 1: Oracle Excess Risk Bound in Terms of Target and Nuisance Errors
  • Theorem 1: Excess Risk Bound for \ref{['alg:gen_est']}
  • Theorem 2: Excess Risk Bound for \ref{['alg:rkhs_est']}
  • Theorem 3: Data Fusion Improves the Excess Risk Upper Bound
  • Lemma 2: Minimax Lower Bound without Data Fusion
  • Theorem 4: Sufficient Conditions for Data Fusion to Improve Worst-Case Performance
  • Lemma S1: Stochastically approximated loss
  • proof : Proof of \ref{['lem:onestep_loss']}
  • Lemma S2: Closed-form solution of kernel ridge regression
  • proof : Proof of \ref{['lem:rkhs_alg']}
  • ...and 10 more