Table of Contents
Fetching ...

Design Stability in Adaptive Experiments: Implications for Treatment Effect Estimation

Saikat Sengupta, Koulik Khamaru, Suvrojit Ghosh, Tirthankar Dasgupta

TL;DR

This paper develops a design-based framework for inference on the average treatment effect under sequential adaptive Bernoulli treatment assignment in finite populations. It proves central limit theorems for inverse propensity weighted (IPW) and augmented IPW (AIPW) estimators under two notions of design stability (strong and weak) and provides conservative variance estimators for valid confidence intervals. The theory is specialized to Wei's adaptive coin design (strong stability, $p^*=\tfrac{1}{2}$) and Efron's biased coin design (weak stability), with empirical results showing that AIPW offers more efficient inference and shorter CIs, especially in non-additive settings. These results enable robust, model-free ATE inference in adaptive, sequential experiments, including online A/B testing and adaptive clinical trials, while clarifying when conservative variance estimation remains valid. The work also connects adaptive allocation dynamics to a stationary imbalance distribution, enriching the theoretical understanding of sequential experimentation in finite populations.

Abstract

We study the problem of estimating the average treatment effect (ATE) under sequentially adaptive treatment assignment mechanisms. In contrast to classical completely randomized designs, we consider a setting in which the probability of assigning treatment to each experimental unit may depend on prior assignments and observed outcomes. Within the potential outcomes framework, we propose and analyze two natural estimators for the ATE: the inverse propensity weighted (IPW) estimator and an augmented IPW (AIPW) estimator. The cornerstone of our analysis is the concept of design stability, which requires that as the number of units grows, either the assignment probabilities converge, or sample averages of the inverse propensity scores and of the inverse complement propensity scores converge in probability to fixed, non-random limits. Our main results establish central limit theorems for both the IPW and AIPW estimators under design stability and provide explicit expressions for their asymptotic variances. We further propose estimators for these variances, enabling the construction of asymptotically valid confidence intervals. Finally, we illustrate our theoretical results in the context of Wei's adaptive coin design and Efron's biased coin design, highlighting the applicability of the proposed methods to sequential experimentation with adaptive randomization.

Design Stability in Adaptive Experiments: Implications for Treatment Effect Estimation

TL;DR

This paper develops a design-based framework for inference on the average treatment effect under sequential adaptive Bernoulli treatment assignment in finite populations. It proves central limit theorems for inverse propensity weighted (IPW) and augmented IPW (AIPW) estimators under two notions of design stability (strong and weak) and provides conservative variance estimators for valid confidence intervals. The theory is specialized to Wei's adaptive coin design (strong stability, ) and Efron's biased coin design (weak stability), with empirical results showing that AIPW offers more efficient inference and shorter CIs, especially in non-additive settings. These results enable robust, model-free ATE inference in adaptive, sequential experiments, including online A/B testing and adaptive clinical trials, while clarifying when conservative variance estimation remains valid. The work also connects adaptive allocation dynamics to a stationary imbalance distribution, enriching the theoretical understanding of sequential experimentation in finite populations.

Abstract

We study the problem of estimating the average treatment effect (ATE) under sequentially adaptive treatment assignment mechanisms. In contrast to classical completely randomized designs, we consider a setting in which the probability of assigning treatment to each experimental unit may depend on prior assignments and observed outcomes. Within the potential outcomes framework, we propose and analyze two natural estimators for the ATE: the inverse propensity weighted (IPW) estimator and an augmented IPW (AIPW) estimator. The cornerstone of our analysis is the concept of design stability, which requires that as the number of units grows, either the assignment probabilities converge, or sample averages of the inverse propensity scores and of the inverse complement propensity scores converge in probability to fixed, non-random limits. Our main results establish central limit theorems for both the IPW and AIPW estimators under design stability and provide explicit expressions for their asymptotic variances. We further propose estimators for these variances, enabling the construction of asymptotically valid confidence intervals. Finally, we illustrate our theoretical results in the context of Wei's adaptive coin design and Efron's biased coin design, highlighting the applicability of the proposed methods to sequential experimentation with adaptive randomization.
Paper Structure (22 sections, 17 theorems, 172 equations, 4 figures)

This paper contains 22 sections, 17 theorems, 172 equations, 4 figures.

Key Result

Theorem 1

Suppose Assumption assn:IPW holds, and the sequential design with inclusion probabilities $\{p_i\}_{i \geq 1}$ is either strongly or weakly stable in the sense of Definition def:design-stability or Definition def:weak-design-stability, respectively. Then the IPW estimator eqn:IPW satisfies with asymptotic variance

Figures (4)

  • Figure 1: Comparison of the theoretical and empirical coverages for Wei's design.
  • Figure 2: Comparison of the average lengths of confidence intervals for Wei's design.
  • Figure 3: Comparison of the theoretical and empirical coverages for Efron's design.
  • Figure 4: Comparison of the average lengths of confidence intervals for Efron's design.

Theorems & Definitions (35)

  • Definition 1: Generalized treatment effect homogeneity
  • Definition 2: Strong design stability
  • Definition 3: Weak design stability
  • Theorem 1
  • Remark 1
  • Theorem 2
  • Remark 2
  • Theorem 3
  • Remark 3
  • Theorem 4
  • ...and 25 more