Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
Ángela López-Cardona, Sebastián Idesis, Mireia Masias-Bruns, Sergi Abadal, Ioannis Arapakis
TL;DR
This paper synthesizes 25 recent fMRI studies (2023–2025) to evaluate two theoretical questions about brain–language model alignment: whether models converge toward a Platonic representation of reality and whether intermediate layers harbor the most brain-like, generalizable features. Using an encoding-model framework across modalities, the authors organize evidence around model scale, task expansion, cross-modal training, and brain-informed tuning, highlighting how these factors influence alignment with neural data. The review finds converging support for both hypotheses: larger, multi-task, and multimodal models tend to align better with brain activity, and intermediate layers often show the strongest correspondences, while final layers are less predictive. However, results are heterogeneous and methodologically varied, suggesting that alignment is robust but not reducible to a single rule, and pointing to cross-modal training as a particularly promising avenue for future research.
Abstract
Do brains and language models converge toward the same internal representations of the world? Recent years have seen a rise in studies of neural activations and model alignment. In this work, we review 25 fMRI-based studies published between 2023 and 2025 and explicitly confront their findings with two key hypotheses: (i) the Platonic Representation Hypothesis -- that as models scale and improve, they converge to a representation of the real world, and (ii) the Intermediate-Layer Advantage -- that intermediate (mid-depth) layers often encode richer, more generalizable features. Our findings provide converging evidence that models and brains may share abstract representational structures, supporting both hypotheses and motivating further research on brain-model alignment.
