MEG-GPT: A transformer-based foundation model for magnetoencephalography data
Rukuang Huang, Sungjun Cho, Chetan Gohil, Oiwi Parker Jones, Mark Woolrich
TL;DR
MEG-GPT presents a foundation-model approach to magnetoencephalography by coupling a data-driven, lossless tokeniser with a transformer-based autoregressive model trained on large-scale resting-state MEG data. The tokeniser enables discrete, high-resolution representation of continuous MEG signals, while the MEG-GPT model learns complex spatiotemporal dependencies and subject-specific patterns through time-attention and next-time-point prediction. The framework demonstrates realistic synthetic MEG generation, captures bursting dynamics and inter-subject variability, and delivers substantial gains in zero-shot decoding across sessions and subjects; fine-tuning further improves performance on new datasets. Together, these results establish MEG-GPT as a scalable foundation model for electrophysiological data with broad implications for neural decoding and computational neuroscience.
Abstract
Modelling the complex spatiotemporal patterns of large-scale brain dynamics is crucial for neuroscience, but traditional methods fail to capture the rich structure in modalities such as magnetoencephalography (MEG). Recent advances in deep learning have enabled significant progress in other domains, such as language and vision, by using foundation models at scale. Here, we introduce MEG-GPT, a transformer based foundation model that uses time-attention and next time-point prediction. To facilitate this, we also introduce a novel data-driven tokeniser for continuous MEG data, which preserves the high temporal resolution of continuous MEG signals without lossy transformations. We trained MEG-GPT on tokenised brain region time-courses extracted from a large-scale MEG dataset (N=612, eyes-closed rest, Cam-CAN data), and show that the learnt model can generate data with realistic spatio-spectral properties, including transient events and population variability. Critically, it performs well in downstream decoding tasks, improving downstream supervised prediction task, showing improved zero-shot generalisation across sessions (improving accuracy from 0.54 to 0.59) and subjects (improving accuracy from 0.41 to 0.49) compared to a baseline methods. Furthermore, we show the model can be efficiently fine-tuned on a smaller labelled dataset to boost performance in cross-subject decoding scenarios. This work establishes a powerful foundation model for electrophysiological data, paving the way for applications in computational neuroscience and neural decoding.
