Table of Contents
Fetching ...

A Relative Error-Based Evaluation Framework of Heterogeneous Treatment Effect Estimators

Jiayi Guo, Haoxuan Li, Ye Tian, Peng Wu

TL;DR

This paper tackles the challenge of evaluating heterogeneous treatment effect estimators without relying on accurate ground-truth counterfactuals. It introduces a robust relative-error framework, $\delta(\hat{\tau}_1,\hat{\tau}_2)$, and derives key nuisance-parameter conditions, enabling $\sqrt{n}$-consistent and asymptotically normal estimation even when outcome models are misspecified, provided the propensity score model is sufficiently consistent. A Dragonnet-inspired neural network with a novel loss $\mathcal{L}_{\text{wls}}$ and balance-regularized constraints, plus a soft-margin approach, stabilizes nuisance estimation and yields valid confidence intervals for relative error. The framework also extends to enhanced CATE estimation via aggregation over pairs of candidate estimators, producing a more robust CATE predictor. Empirical results on IHDP, Twins, and Jobs demonstrate reliable relative-error inference, improved estimator selection, and competitive or superior CATE performance, highlighting practical impact for reliable HTE evaluation and learning in real-world settings.

Abstract

While significant progress has been made in heterogeneous treatment effect (HTE) estimation, the evaluation of HTE estimators remains underdeveloped. In this article, we propose a robust evaluation framework based on relative error, which quantifies performance differences between two HTE estimators. We first derive the key theoretical conditions on the nuisance parameters that are necessary to achieve a robust estimator of relative error. Building on these conditions, we introduce novel loss functions and design a neural network architecture to estimate nuisance parameters and obtain robust estimation of relative error, thereby achieving reliable evaluation of HTE estimators. We provide the large sample properties of the proposed relative error estimator. Furthermore, beyond evaluation, we propose a new learning algorithm for HTE that leverages both the previously HTE estimators and the nuisance parameters learned through our neural network architecture. Extensive experiments demonstrate that our evaluation framework supports reliable comparisons across HTE estimators, and the proposed learning algorithm for HTE exhibits desirable performance.

A Relative Error-Based Evaluation Framework of Heterogeneous Treatment Effect Estimators

TL;DR

This paper tackles the challenge of evaluating heterogeneous treatment effect estimators without relying on accurate ground-truth counterfactuals. It introduces a robust relative-error framework, , and derives key nuisance-parameter conditions, enabling -consistent and asymptotically normal estimation even when outcome models are misspecified, provided the propensity score model is sufficiently consistent. A Dragonnet-inspired neural network with a novel loss and balance-regularized constraints, plus a soft-margin approach, stabilizes nuisance estimation and yields valid confidence intervals for relative error. The framework also extends to enhanced CATE estimation via aggregation over pairs of candidate estimators, producing a more robust CATE predictor. Empirical results on IHDP, Twins, and Jobs demonstrate reliable relative-error inference, improved estimator selection, and competitive or superior CATE performance, highlighting practical impact for reliable HTE evaluation and learning in real-world settings.

Abstract

While significant progress has been made in heterogeneous treatment effect (HTE) estimation, the evaluation of HTE estimators remains underdeveloped. In this article, we propose a robust evaluation framework based on relative error, which quantifies performance differences between two HTE estimators. We first derive the key theoretical conditions on the nuisance parameters that are necessary to achieve a robust estimator of relative error. Building on these conditions, we introduce novel loss functions and design a neural network architecture to estimate nuisance parameters and obtain robust estimation of relative error, thereby achieving reliable evaluation of HTE estimators. We provide the large sample properties of the proposed relative error estimator. Furthermore, beyond evaluation, we propose a new learning algorithm for HTE that leverages both the previously HTE estimators and the nuisance parameters learned through our neural network architecture. Extensive experiments demonstrate that our evaluation framework supports reliable comparisons across HTE estimators, and the proposed learning algorithm for HTE exhibits desirable performance.
Paper Structure (26 sections, 2 theorems, 41 equations, 3 figures, 10 tables)

This paper contains 26 sections, 2 theorems, 41 equations, 3 figures, 10 tables.

Key Result

Theorem 1

If the propensity score model is correctly specified, and $\check \gamma$, $\check \beta_0$ as well as $\check \beta_1$ converge to their probability limits at a rate faster than $n^{-1/4}$, then we have where $\sigma^{2} = \text{Var}\{\varphi(Z; \bar{u}_{0}, \bar{u}_{1}, \bar{e} )\}$ and $\xrightarrow{d}$ means convergence in distribution.

Figures (3)

  • Figure 1: Coverage rate on IHDP and Twins.
  • Figure 2: Selection accuracy on IHDP and Twins.
  • Figure 3: Neural Network Structure

Theorems & Definitions (4)

  • Example 1: A misspecified model
  • Theorem 1
  • Proposition 2
  • proof : Proof of Proposition 2