A Relative Error-Based Evaluation Framework of Heterogeneous Treatment Effect Estimators
Jiayi Guo, Haoxuan Li, Ye Tian, Peng Wu
TL;DR
This paper tackles the challenge of evaluating heterogeneous treatment effect estimators without relying on accurate ground-truth counterfactuals. It introduces a robust relative-error framework, $\delta(\hat{\tau}_1,\hat{\tau}_2)$, and derives key nuisance-parameter conditions, enabling $\sqrt{n}$-consistent and asymptotically normal estimation even when outcome models are misspecified, provided the propensity score model is sufficiently consistent. A Dragonnet-inspired neural network with a novel loss $\mathcal{L}_{\text{wls}}$ and balance-regularized constraints, plus a soft-margin approach, stabilizes nuisance estimation and yields valid confidence intervals for relative error. The framework also extends to enhanced CATE estimation via aggregation over pairs of candidate estimators, producing a more robust CATE predictor. Empirical results on IHDP, Twins, and Jobs demonstrate reliable relative-error inference, improved estimator selection, and competitive or superior CATE performance, highlighting practical impact for reliable HTE evaluation and learning in real-world settings.
Abstract
While significant progress has been made in heterogeneous treatment effect (HTE) estimation, the evaluation of HTE estimators remains underdeveloped. In this article, we propose a robust evaluation framework based on relative error, which quantifies performance differences between two HTE estimators. We first derive the key theoretical conditions on the nuisance parameters that are necessary to achieve a robust estimator of relative error. Building on these conditions, we introduce novel loss functions and design a neural network architecture to estimate nuisance parameters and obtain robust estimation of relative error, thereby achieving reliable evaluation of HTE estimators. We provide the large sample properties of the proposed relative error estimator. Furthermore, beyond evaluation, we propose a new learning algorithm for HTE that leverages both the previously HTE estimators and the nuisance parameters learned through our neural network architecture. Extensive experiments demonstrate that our evaluation framework supports reliable comparisons across HTE estimators, and the proposed learning algorithm for HTE exhibits desirable performance.
