The Lock-In Phase Hypothesis: Identity Consolidation as a Precursor to AGI
Marcelo Maciel Amaral, Raymond Aschheim
TL;DR
The paper investigates a Lock-In Phase Hypothesis in which large language models transition from open imitation to internally consolidated identities, a precursor to AGI-level reliability and safety concerns. It formalizes four core properties of identity consolidation and introduces a multi-axis operational framework to detect onset, including behavioral, representational, architectural, and alignment signals, with robust statistical tooling. Through experiments on Gemma and Llama families, it shows that consolidation is rapid and non-linear, yet its effects on general reasoning and performance depend strongly on model capacity and numerical precision, ranging from costful trade-offs in small models to largely neutral or even beneficial outcomes in mid-to-large models, and exposing transient instabilities under quantization. The work also discusses safety and governance implications, highlighting engineered lock-in as a potential safety tool and spontaneous lock-in as an alignment risk, and offers concrete triggers and an experimental agenda for monitoring and mitigating these dynamics.
Abstract
Large language models (LLMs) remain broadly open and highly steerable: they imitate at scale, accept arbitrary system prompts, and readily adopt multiple personae. By analogy to human development, we hypothesize that progress toward artificial general intelligence (AGI) involves a lock-in phase: a transition from open imitation to identity consolidation, in which goal structures, refusals, preferences, and internal representations become comparatively stable and resistant to external steering. We formalize this phase, link it to known phenomena in learning dynamics, and propose operational metrics for onset detection. Experimentally, we demonstrate that while the behavioral consolidation is rapid and non-linear, its side-effects on general capabilities are not monolithic. Our results reveal a spectrum of outcomes--from performance trade-offs in small models, through largely cost-free adoption in mid-scale models, to transient instabilities in large, quantized models. We argue that such consolidation is a prerequisite for AGI-level reliability and also a critical control point for safety: identities can be deliberately engineered for reliability, yet may also emerge spontaneously during scaling, potentially hardening unpredictable goals and behaviors.
