Table of Contents
Fetching ...

Transfer Learning Beyond the Standard Model

Veena Krishnaraj, Adrian E. Bayer, Christian Kragh Jespersen, Peter Melchior

TL;DR

This work investigates whether transfer learning from $\Lambda$CDM can enable inference of beyond-$\Lambda$CDM physics using significantly fewer simulations. A two-stage framework with a dummy bottleneck is used to pre-train on $\Lambda$CDM and fine-tune on massive neutrinos, modified gravity, and primordial non-Gaussianities, employing both the power spectrum $P(k)$ and the marked power spectrum $MP(k)$ from Quijote simulations. The study finds that the dummy-bottleneck approach generally reduces the required beyond-$\Lambda$CDM simulations and improves total MSE, though strong degeneracies such as between $\sigma_8$ and $M_\nu$ can cause negative transfer, especially with the marked summary. Gains depend on degeneracies, data summaries, and pre-training size (e.g., ~2k vs ~22k simulations). The results highlight both the promise and pitfalls of foundation-model-like pretraining for physics, with implications for applying similar strategies to other observables and posterior inference while cautioning about potential biases from pretraining on standard-model priors.

Abstract

Machine learning enables powerful cosmological inference but typically requires many high-fidelity simulations covering many cosmological models. Transfer learning offers a way to reduce the simulation cost by reusing knowledge across models. We show that pre-training on the standard model of cosmology, $Λ$CDM, and fine-tuning on various beyond-$Λ$CDM scenarios -- including massive neutrinos, modified gravity, and primordial non-Gaussianities -- can enable inference with significantly fewer beyond-$Λ$CDM simulations. However, we also show that negative transfer can occur when strong physical degeneracies exist between $Λ$CDM and beyond-$Λ$CDM parameters. We consider various transfer architectures, finding that including bottleneck structures provides the best performance. Our findings illustrate the opportunities and pitfalls of foundation-model approaches in physics: pre-training can accelerate inference, but may also hinder learning new physics.

Transfer Learning Beyond the Standard Model

TL;DR

This work investigates whether transfer learning from CDM can enable inference of beyond-CDM physics using significantly fewer simulations. A two-stage framework with a dummy bottleneck is used to pre-train on CDM and fine-tune on massive neutrinos, modified gravity, and primordial non-Gaussianities, employing both the power spectrum and the marked power spectrum from Quijote simulations. The study finds that the dummy-bottleneck approach generally reduces the required beyond-CDM simulations and improves total MSE, though strong degeneracies such as between and can cause negative transfer, especially with the marked summary. Gains depend on degeneracies, data summaries, and pre-training size (e.g., ~2k vs ~22k simulations). The results highlight both the promise and pitfalls of foundation-model-like pretraining for physics, with implications for applying similar strategies to other observables and posterior inference while cautioning about potential biases from pretraining on standard-model priors.

Abstract

Machine learning enables powerful cosmological inference but typically requires many high-fidelity simulations covering many cosmological models. Transfer learning offers a way to reduce the simulation cost by reusing knowledge across models. We show that pre-training on the standard model of cosmology, CDM, and fine-tuning on various beyond-CDM scenarios -- including massive neutrinos, modified gravity, and primordial non-Gaussianities -- can enable inference with significantly fewer beyond-CDM simulations. However, we also show that negative transfer can occur when strong physical degeneracies exist between CDM and beyond-CDM parameters. We consider various transfer architectures, finding that including bottleneck structures provides the best performance. Our findings illustrate the opportunities and pitfalls of foundation-model approaches in physics: pre-training can accelerate inference, but may also hinder learning new physics.
Paper Structure (9 sections, 7 figures)

This paper contains 9 sections, 7 figures.

Figures (7)

  • Figure 1: Dummy network architecture. The model takes the (marked) power spectrum $P(k)$ as input and outputs cosmological parameters $\theta_{\Lambda{\rm CDM}}$. Additional latent "dummy" nodes $\psi_{\rm dummy}$ are included in the output layer to provide extra representational capacity for fine-tuning.
  • Figure 2: Test MSE as a function of the number of fine-tuning simulations for the massive neutrino cosmology using standard (top) and marked (bottom) power spectra for $\sigma_8$ (left), $M_\nu$ (center), and the total MSE across all normalized parameters (right). Transfer learning using a dummy node (red) always outperforms the result with no transfer learning (black) in terms of the total MSE, however negative transfer occurs for the marked power spectrum for $\sigma_8$ and $M_\nu$ due to the physical degeneracy between $M_\nu$ and $\sigma_8$. Other transfer learning architectures (teal, yellow) are suboptimal and result in more severe negative transfer.
  • Figure 3: Total MSE across all normalized parameters for modified gravity (left), equilateral (center), and local (right) primordial non-Gaussianity cosmologies. The colored lines represent different pre-training set sizes, which outperform the model trained directly on beyond-$\Lambda$CDM without transfer learning (black), except in the case of local $f_{\rm NL}$ due to the prior.
  • Figure 4: Extension of Figure \ref{['fig:mse_nwLH']} showing the MSE for all individual parameters in the massive neutrino cosmology, but only for the dummy node architecture. Transfer learning provides improvements for some $\Lambda$CDM parameters when training data is very limited and when using the power spectrum (left), but offers little to no benefit for $\sigma_8$ and $M_\nu$. In fact, for the marked power spectrum (right) it can even degrade performance (negative transfer) at low simulation counts.
  • Figure 5: Same as Figure \ref{['fig:mse_beyond_lcdm']}, but showing MSE for each individual parameter in the modified gravity, equilateral, and local non-Gaussianity cosmologies.
  • ...and 2 more figures