Table of Contents
Fetching ...

A Definition of AGI

Dan Hendrycks, Dawn Song, Christian Szegedy, Honglak Lee, Yarin Gal, Erik Brynjolfsson, Sharon Li, Andy Zou, Lionel Levine, Bo Han, Jie Fu, Ziwei Liu, Jinwoo Shin, Kimin Lee, Mantas Mazeika, Long Phan, George Ingebretsen, Adam Khoja, Cihang Xie, Olawale Salaudeen, Matthias Hein, Kevin Zhao, Alexander Pan, David Duvenaud, Bo Li, Steve Omohundro, Gabriel Alfour, Max Tegmark, Kevin McGrew, Gary Marcus, Jaan Tallinn, Eric Schmidt, Yoshua Bengio

TL;DR

<3-5 sentence high-level summary> The paper establishes a quantifiable AGI definition anchored in CHC theory and operationalizes it through ten equal-weight cognitive domains tested with adapted human psychometric batteries across modalities. It reports that contemporary models exhibit a highly uneven, or 'jagged', cognitive profile, with strong performance in knowledge-related tasks but critical deficits in long-term memory storage and other foundational mechanisms, resulting in an estimated AGI score of about 27% for GPT-4 and 57% for GPT-5. The framework provides precise diagnostics of strengths and bottlenecks, illustrating how retrieval, memory, and multimodal integration constrain progress toward true human-level AGI. It also discusses methodological considerations, limitations, and the risk of relying on aggregate scores that can mask severe bottlenecks, underscoring the need for continual learning and memory systems beyond large-scale pretraining.</p>

Abstract

The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quantifiable framework to address this, defining AGI as matching the cognitive versatility and proficiency of a well-educated adult. To operationalize this, we ground our methodology in Cattell-Horn-Carroll theory, the most empirically validated model of human cognition. The framework dissects general intelligence into ten core cognitive domains-including reasoning, memory, and perception-and adapts established human psychometric batteries to evaluate AI systems. Application of this framework reveals a highly "jagged" cognitive profile in contemporary models. While proficient in knowledge-intensive domains, current AI systems have critical deficits in foundational cognitive machinery, particularly long-term memory storage. The resulting AGI scores (e.g., GPT-4 at 27%, GPT-5 at 57%) concretely quantify both rapid progress and the substantial gap remaining before AGI.

A Definition of AGI

TL;DR

<3-5 sentence high-level summary> The paper establishes a quantifiable AGI definition anchored in CHC theory and operationalizes it through ten equal-weight cognitive domains tested with adapted human psychometric batteries across modalities. It reports that contemporary models exhibit a highly uneven, or 'jagged', cognitive profile, with strong performance in knowledge-related tasks but critical deficits in long-term memory storage and other foundational mechanisms, resulting in an estimated AGI score of about 27% for GPT-4 and 57% for GPT-5. The framework provides precise diagnostics of strengths and bottlenecks, illustrating how retrieval, memory, and multimodal integration constrain progress toward true human-level AGI. It also discusses methodological considerations, limitations, and the risk of relying on aggregate scores that can mask severe bottlenecks, underscoring the need for continual learning and memory systems beyond large-scale pretraining.</p>

Abstract

The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quantifiable framework to address this, defining AGI as matching the cognitive versatility and proficiency of a well-educated adult. To operationalize this, we ground our methodology in Cattell-Horn-Carroll theory, the most empirically validated model of human cognition. The framework dissects general intelligence into ten core cognitive domains-including reasoning, memory, and perception-and adapts established human psychometric batteries to evaluate AI systems. Application of this framework reveals a highly "jagged" cognitive profile in contemporary models. While proficient in knowledge-intensive domains, current AI systems have critical deficits in foundational cognitive machinery, particularly long-term memory storage. The resulting AGI scores (e.g., GPT-4 at 27%, GPT-5 at 57%) concretely quantify both rapid progress and the substantial gap remaining before AGI.
Paper Structure (129 sections, 13 figures, 1 table)

This paper contains 129 sections, 13 figures, 1 table.

Figures (13)

  • Figure 1: The capabilities of GPT-4 and GPT-5. Here GPT-5 answers questions in 'Auto' mode.
  • Figure 2: The ten core cognitive components of our AGI definition.
  • Figure 3: Intelligence as a processor. Figure based on McGrew_Schneider_2018_CHCTheoryRevised.
  • Figure :
  • Figure :
  • ...and 8 more figures