Transfer Learning Beyond the Standard Model
Veena Krishnaraj, Adrian E. Bayer, Christian Kragh Jespersen, Peter Melchior
TL;DR
This work investigates whether transfer learning from $\Lambda$CDM can enable inference of beyond-$\Lambda$CDM physics using significantly fewer simulations. A two-stage framework with a dummy bottleneck is used to pre-train on $\Lambda$CDM and fine-tune on massive neutrinos, modified gravity, and primordial non-Gaussianities, employing both the power spectrum $P(k)$ and the marked power spectrum $MP(k)$ from Quijote simulations. The study finds that the dummy-bottleneck approach generally reduces the required beyond-$\Lambda$CDM simulations and improves total MSE, though strong degeneracies such as between $\sigma_8$ and $M_\nu$ can cause negative transfer, especially with the marked summary. Gains depend on degeneracies, data summaries, and pre-training size (e.g., ~2k vs ~22k simulations). The results highlight both the promise and pitfalls of foundation-model-like pretraining for physics, with implications for applying similar strategies to other observables and posterior inference while cautioning about potential biases from pretraining on standard-model priors.
Abstract
Machine learning enables powerful cosmological inference but typically requires many high-fidelity simulations covering many cosmological models. Transfer learning offers a way to reduce the simulation cost by reusing knowledge across models. We show that pre-training on the standard model of cosmology, $Λ$CDM, and fine-tuning on various beyond-$Λ$CDM scenarios -- including massive neutrinos, modified gravity, and primordial non-Gaussianities -- can enable inference with significantly fewer beyond-$Λ$CDM simulations. However, we also show that negative transfer can occur when strong physical degeneracies exist between $Λ$CDM and beyond-$Λ$CDM parameters. We consider various transfer architectures, finding that including bottleneck structures provides the best performance. Our findings illustrate the opportunities and pitfalls of foundation-model approaches in physics: pre-training can accelerate inference, but may also hinder learning new physics.
