Table of Contents
Fetching ...

A Coherence-Based Measure of AGI

Fares Fourati

TL;DR

The paper argues that arithmetic averaging of CHC-domain scores overstates AI generality by permitting compensability. It introduces a coherence-based framework using generalized means AGI_p across a compensability parameter p and aggregates these into AGI_AUC to quantify robustness to stricter non-compensatory regimes. Applied to CHC-domain scores for GPT-4/5 and to a 17-benchmark suite, the approach reveals persistent bottlenecks and imbalances masked by arithmetic means, providing a stricter, interpretable measure of progress toward general intelligence. The results suggest that true generality requires balanced, cross-domain competence, and propose coherence-aware reporting alongside traditional aggregates to better guide benchmark design and development toward genuinely general systems.

Abstract

Recent approaches to evaluating Artificial General Intelligence (AGI) typically summarize a system's capability using the arithmetic mean of its proficiencies across multiple cognitive domains. While simple, this implicitly assumes compensability: exceptional performance in some areas can offset severe deficiencies in others. Genuine general intelligence, however, requires coherent sufficiency: balanced competence across all essential faculties. We introduce a coherence-based measure of AGI that integrates the generalized mean over a continuum of compensability exponents. This yields an area-under-the-curve (AUC) metric spanning arithmetic, geometric, and harmonic regimes, quantifying how robust an evaluated capability remains as compensability assumptions become stricter. Unlike the arithmetic mean, which rewards specialization, the AUC penalizes imbalance and exposes bottlenecks that constrain performance. To illustrate the framework, we apply it to cognitive profiles derived from the Cattell-Horn-Carroll (CHC) model, showing how coherence-based aggregation highlights imbalances that are obscured by arithmetic averaging. As a second, independent example, we apply the same methodology to a set of 17 heterogeneous benchmarks, demonstrating how coherence-based evaluation can reveal unevenness even in narrower task collections. These examples show that the proposed approach offers a principled, interpretable, and stricter foundation for measuring progress toward AGI.

A Coherence-Based Measure of AGI

TL;DR

The paper argues that arithmetic averaging of CHC-domain scores overstates AI generality by permitting compensability. It introduces a coherence-based framework using generalized means AGI_p across a compensability parameter p and aggregates these into AGI_AUC to quantify robustness to stricter non-compensatory regimes. Applied to CHC-domain scores for GPT-4/5 and to a 17-benchmark suite, the approach reveals persistent bottlenecks and imbalances masked by arithmetic means, providing a stricter, interpretable measure of progress toward general intelligence. The results suggest that true generality requires balanced, cross-domain competence, and propose coherence-aware reporting alongside traditional aggregates to better guide benchmark design and development toward genuinely general systems.

Abstract

Recent approaches to evaluating Artificial General Intelligence (AGI) typically summarize a system's capability using the arithmetic mean of its proficiencies across multiple cognitive domains. While simple, this implicitly assumes compensability: exceptional performance in some areas can offset severe deficiencies in others. Genuine general intelligence, however, requires coherent sufficiency: balanced competence across all essential faculties. We introduce a coherence-based measure of AGI that integrates the generalized mean over a continuum of compensability exponents. This yields an area-under-the-curve (AUC) metric spanning arithmetic, geometric, and harmonic regimes, quantifying how robust an evaluated capability remains as compensability assumptions become stricter. Unlike the arithmetic mean, which rewards specialization, the AUC penalizes imbalance and exposes bottlenecks that constrain performance. To illustrate the framework, we apply it to cognitive profiles derived from the Cattell-Horn-Carroll (CHC) model, showing how coherence-based aggregation highlights imbalances that are obscured by arithmetic averaging. As a second, independent example, we apply the same methodology to a set of 17 heterogeneous benchmarks, demonstrating how coherence-based evaluation can reveal unevenness even in narrower task collections. These examples show that the proposed approach offers a principled, interpretable, and stricter foundation for measuring progress toward AGI.
Paper Structure (36 sections, 6 equations, 2 figures, 13 tables)

This paper contains 36 sections, 6 equations, 2 figures, 13 tables.

Figures (2)

  • Figure 1: Comparison of model performance across aggregation exponents $p$. Curves show $\mathrm{AGI}_p$ values; shaded regions indicate the AUC.
  • Figure 2: Comparison of model performance across aggregation exponents $p$. Curves show $\mathrm{AGI}_p$ values derived from the 17 benchmarks in Table \ref{['tab:benchmark_comparison']}.