Assessing the robustness of heterogeneous treatment effects in survival analysis under informative censoring
Yuxin Wang, Dennis Frauen, Jonas Schweisthal, Maresa Schröder, Stefan Feuerriegel
TL;DR
This work tackles the challenge of estimating heterogeneous treatment effects in survival data when censoring is informative. By reframing the problem with partial identification, the authors derive informative lower and upper bounds on the conditional average treatment effect (CATE) and its component survival quantities, incorporating censoring strength and a domain-knowledge post-dropout survival function. They introduce SurvB-learner, a model-agnostic two-stage meta-learner that produces consistent, double-robust, and quasi-oracle-efficient estimates of these bounds, applicable to discrete and continuous treatments and adaptable to real-world data with hidden confounding. Through synthetic experiments and an application to the ADJUVANT gefitinib trial in NSCLC, the framework reveals genomically defined subgroups with robust treatment benefits despite dropout, illustrating its practical value for robust, individualized evidence generation in medicine and epidemiology.
Abstract
Dropout is common in clinical studies, with up to half of patients leaving early due to side effects or other reasons. When dropout is informative (i.e., dependent on survival time), it introduces censoring bias, because of which treatment effect estimates are also biased. In this paper, we propose an assumption-lean framework to assess the robustness of conditional average treatment effect (CATE) estimates in survival analysis when facing censoring bias. Unlike existing works that rely on strong assumptions, such as non-informative censoring, to obtain point estimation, we use partial identification to derive informative bounds on the CATE. Thereby, our framework helps to identify patient subgroups where treatment is effective despite informative censoring. We further develop a novel meta-learner that estimates the bounds using arbitrary machine learning models and with favorable theoretical properties, including double robustness and quasi-oracle efficiency. We demonstrate the practical value of our meta-learner through numerical experiments and in an application to a cancer drug trial. Together, our framework offers a practical tool for assessing the robustness of estimated treatment effects in the presence of censoring and thus promotes the reliable use of survival data for evidence generation in medicine and epidemiology.
