A Simple Geometric Proof of the Optimality of the Sequential Probability Ratio Test for Symmetric Bernoulli Hypotheses
Chirag Pabbaraju, Gregory Valiant, Rishi Verma
TL;DR
The paper provides a self-contained, constructive proof of the optimality of the Sequential Probability Ratio Test for distinguishing $p=\tfrac{1}{2}\pm\varepsilon$ by modeling coin flips as a biased random walk on a 2D grid and reinterpreting strategies as lattice colorings. It introduces a greedy, local-move framework that converts any policy into a linear difference policy $\mathcal{P}_c$, and shows that for each tradeoff parameter $\beta>0$ there exists a corresponding $c(\beta)$ that minimizes the Bayes-like risk $R(\beta)=(\delta^++\delta^-)+\beta(H^++H^-)$. The argument proceeds through five concrete steps—triangular recoloring, truncation to a linear frontier, gap-filling, level-extensions, and erasure of extraneous points—demonstrating that the linear policy achieves at least as good a risk as any alternative. The results recover the SPRT’s optimality in this symmetric Bernoulli setting and are presented in a way that invites generalization to broader Bayes risks and non-symmetric hypotheses.
Abstract
This paper revisits the classical problem of determining the bias of a weighted coin, where the bias is known to be either $p = 1/2 + \varepsilon$ or $p = 1/2 - \varepsilon$, while minimizing the expected number of coin tosses and the error probability. The optimal strategy for this problem is given by Wald's Sequential Probability Ratio Test (SPRT), which compares the log-likelihood ratio against fixed thresholds to determine a stopping time. Classical proofs of this result typically rely on analytical, continuous, and non-constructive arguments. In this paper, we present a discrete, self-contained proof of the optimality of the SPRT for this problem. We model the problem as a biased random walk on the two-dimensional (heads, tails) integer lattice, and model strategies as marked stopping times on this lattice. Our proof takes a straightforward greedy approach, showing how any arbitrary strategy may be transformed into the optimal, parallel-line "difference policy" corresponding to the SPRT, via a sequence of local perturbations that improve a Bayes risk objective.
