Knowledge Divergence and the Value of Debate for Scalable Oversight

Robin Young

Knowledge Divergence and the Value of Debate for Scalable Oversight

Robin Young

TL;DR

This work offers the first formal connection between debate and RLAIF, a geometric foundation for understanding when adversarial oversight protocols are justified, and connection to the problem of eliciting latent knowledge across models with complementary information.

Abstract

AI safety via debate and reinforcement learning from AI feedback (RLAIF) are both proposed methods for scalable oversight of advanced AI systems, yet no formal framework relates them or characterizes when debate offers an advantage. We analyze this by parameterizing debate's value through the geometry of knowledge divergence between debating models. Using principal angles between models' representation subspaces, we prove that the debate advantage admits an exact closed form. When models share identical training corpora, debate reduces to RLAIF-like where a single-agent method recovers the same optimum. When models possess divergent knowledge, debate advantage scales with a phase transition from quadratic regime (debate offers negligible benefit) to linear regime (debate is essential). We classify three regimes of knowledge divergence (shared, one-sided, and compositional) and provide existence results showing that debate can achieve outcomes inaccessible to either model alone, alongside a negative result showing that sufficiently strong adversarial incentives cause coordination failure in the compositional regime, with a sharp threshold separating effective from ineffective debate. We offer the first formal connection between debate and RLAIF, a geometric foundation for understanding when adversarial oversight protocols are justified, and connection to the problem of eliciting latent knowledge across models with complementary information.

Knowledge Divergence and the Value of Debate for Scalable Oversight

TL;DR

Abstract

Paper Structure (15 sections, 15 theorems, 25 equations)

This paper contains 15 sections, 15 theorems, 25 equations.

Introduction
Geometric Framework
Setup and Notation
RLAIF and Debate as Optimization
Decomposition via Principal Vectors
Main Result
Corollaries
Multi-Agent Debate
Discrete Regimes of Knowledge Divergence
One-Sided Revelation
Compositional Existence
Limits of Adversarial Debate
Dynamic Subspaces Under Debate
Connection to Constitutional Ambiguity
Discussion

Key Result

Lemma 4

The set $\{u_1, \ldots, u_k, \tilde{v}_1, \ldots, \tilde{v}_m\}$, where $m = |\{i : \theta_i > 0\}|$, forms an orthonormal basis for $V_A + V_B$.

Theorems & Definitions (38)

Definition 1: Principal Angles
Definition 2: Linear Constitutional Scoring
Definition 3: Debate Advantage
Lemma 4
proof
Definition 5: Private Information Value
Theorem 6: Debate Advantage Bound
proof
Remark 1: Scaling regimes
Corollary 7: Same-Corpus Equivalence
...and 28 more

Knowledge Divergence and the Value of Debate for Scalable Oversight

TL;DR

Abstract

Knowledge Divergence and the Value of Debate for Scalable Oversight

Authors

TL;DR

Abstract

Table of Contents

Key Result

Theorems & Definitions (38)