Table of Contents
Fetching ...

From Checklists to Clusters: A Homeostatic Account of AGI Evaluation

Brett Reynolds

TL;DR

The paper critiques current AGI evaluations for using equal domain weights and snapshot scoring, arguing these practices miss how general intelligence is maintained under perturbation. It introduces a homeostatic-property-cluster framework that combines CHC-derived g-loadings with structural priors to form a centrality-weighted evaluation, and pairs this with a stability battery comprising three indices—Profile Stability (pCSI), Durable Learning (dCSI), and Error-Decay (eCSI)—aggregated into CSI to assess robustness. It then provides concrete, black-box evaluation protocols, including perturbation-battery design, pre-registration, and anti-gaming measures, while distinguishing compensatory tool use from contortions that merely inflate scores. The framework yields testable predictions about how durable learning, memory consolidation, and stability relate to real-world deployment success, and discusses governance implications, limitations, and possible extensions to architecture-specific perturbations and cross-species comparisons. Overall, the approach aims to shift AGI evaluation from breadth-only snapshots toward multidomain, stability-aware profiling that better predicts robust, deployable general intelligence.

Abstract

Contemporary AGI evaluations report multidomain capability profiles, yet they typically assign symmetric weights and rely on snapshot scores. This creates two problems: (i) equal weighting treats all domains as equally important when human intelligence research suggests otherwise, and (ii) snapshot testing can't distinguish durable capabilities from brittle performances that collapse under delay or stress. I argue that general intelligence -- in humans and potentially in machines -- is better understood as a homeostatic property cluster: a set of abilities plus the mechanisms that keep those abilities co-present under perturbation. On this view, AGI evaluation should weight domains by their causal centrality (their contribution to cluster stability) and require evidence of persistence across sessions. I propose two battery-compatible extensions: a centrality-prior score that imports CHC-derived weights with transparent sensitivity analysis, and a Cluster Stability Index family that separates profile persistence, durable learning, and error correction. These additions preserve multidomain breadth while reducing brittleness and gaming. I close with testable predictions and black-box protocols labs can adopt without architectural access.

From Checklists to Clusters: A Homeostatic Account of AGI Evaluation

TL;DR

The paper critiques current AGI evaluations for using equal domain weights and snapshot scoring, arguing these practices miss how general intelligence is maintained under perturbation. It introduces a homeostatic-property-cluster framework that combines CHC-derived g-loadings with structural priors to form a centrality-weighted evaluation, and pairs this with a stability battery comprising three indices—Profile Stability (pCSI), Durable Learning (dCSI), and Error-Decay (eCSI)—aggregated into CSI to assess robustness. It then provides concrete, black-box evaluation protocols, including perturbation-battery design, pre-registration, and anti-gaming measures, while distinguishing compensatory tool use from contortions that merely inflate scores. The framework yields testable predictions about how durable learning, memory consolidation, and stability relate to real-world deployment success, and discusses governance implications, limitations, and possible extensions to architecture-specific perturbations and cross-species comparisons. Overall, the approach aims to shift AGI evaluation from breadth-only snapshots toward multidomain, stability-aware profiling that better predicts robust, deployable general intelligence.

Abstract

Contemporary AGI evaluations report multidomain capability profiles, yet they typically assign symmetric weights and rely on snapshot scores. This creates two problems: (i) equal weighting treats all domains as equally important when human intelligence research suggests otherwise, and (ii) snapshot testing can't distinguish durable capabilities from brittle performances that collapse under delay or stress. I argue that general intelligence -- in humans and potentially in machines -- is better understood as a homeostatic property cluster: a set of abilities plus the mechanisms that keep those abilities co-present under perturbation. On this view, AGI evaluation should weight domains by their causal centrality (their contribution to cluster stability) and require evidence of persistence across sessions. I propose two battery-compatible extensions: a centrality-prior score that imports CHC-derived weights with transparent sensitivity analysis, and a Cluster Stability Index family that separates profile persistence, durable learning, and error correction. These additions preserve multidomain breadth while reducing brittleness and gaming. I close with testable predictions and black-box protocols labs can adopt without architectural access.
Paper Structure (86 sections, 14 equations, 3 figures)

This paper contains 86 sections, 14 equations, 3 figures.

Figures (3)

  • Figure 1: Performance trajectories under scaffold degradation. Compensatory scaffolds enable graceful degradation with preserved profile shape (high $\mathrm{pCSI}$). Contorted scaffolds produce catastrophic collapse when removed (low $\mathrm{pCSI}$). Both systems start at 85% with full scaffolds, but only the compensatory system maintains robust capabilities when scaffolds are degraded or removed. Illustrative data.
  • Figure 2: Sensitivity of AGI score to weighting scheme for a model with jagged profile: high on central domains (R: 90%, MS: 85%, MR: 80%), low on peripheral domains (S: 40%, A: 35%). The centrality-prior score ranges from 71--76% as $\lambda$ varies, consistently exceeding the equal-weight score of 68%. The sensitivity band quantifies uncertainty while showing that centrality weighting systematically rewards capability profiles aligned with structural importance. Illustrative data.
  • Figure 3: Multi-dimensional governance space. Systems have to meet both capability breadth (centrality-prior score) and stability ($\mathrm{CSI}$) thresholds for governance tiers. The nested regions show escalating requirements: Tier A requires moderate breadth and stability; Tier B requires both to be high; Tier C requires near-ceiling performance on both dimensions. High score alone is insufficient: a system at (55%, 0.91) has excellent stability but lacks capability breadth and doesn't trigger enhanced governance. The conjunction requirement makes gaming harder and better captures general intelligence. Example systems shown as points; thresholds are illustrative templates subject to empirical calibration.