Table of Contents
Fetching ...

Simulation budgeting for hybrid effective field theories

Alexa Bartlett, Joseph DeRose, Martin White

TL;DR

This work establishes a practical framework for budgeting N-body simulations to train hybrid effective field theory (HEFT) emulators for large-scale structure, balancing perturbation theory with N-body displacements to reach ~1% accuracy in the targeted k-range. By quantifying survey-driven statistical and modeling uncertainties (including intrinsic alignments and baryonic feedback) and testing higher-order bias terms, the authors derive explicit simulation requirements (box size, particle load, starting redshift) and present a tiered emulator training strategy across an 8-parameter cosmology space that includes $w_0w_a$CDM+$\sum m_\nu$. They demonstrate, via surrogate modeling and targeted use cases (DES Y6 GGL, SO CMB lensing, Roman cosmic shear, and steel samples), that emulators with sub-percent precision are achievable with on the order of a few hundred to a few thousand simulations, depending on the parameter volume and $k_{max}$. The results provide concrete budgeting guidelines for efficient HEFT emulator design in future cosmological analyses. The methodology combines Lagrangian perturbation theory, HEFT, variance-reduction techniques (ZCV), and a PCA+PCE emulator to map cosmology-dependent power spectra across scales relevant for 2-point statistics.

Abstract

In this work, we forecast the number of, and requirements on, N-body simulations needed to train hybrid effective field theory (HEFT) emulators for a range of use cases, using a hybrid of HMcode and perturbation theory as a surrogate model. Our accuracy goals, determined with careful consideration of statistical and systematic uncertainties, are $1\%$ accurate in the high-likelihood range of cosmological parameters, and $2\%$ accurate over a broader parameter space volume for $k<1 h Mpc^{-1}$ and $z<3$. Focusing in part on the 8-parameter $w_0w_a$CDM+$m_ν$ cosmological model, we find that $<225$ simulations are required to meet our error goals over our wide parameter space, including models with rapidly evolving dark energy, given our simulation and emulator recommendations. For a more restricted parameter space volume, as few as 80 simulations are sufficient. We additionally present simulation forecasts for example use cases, and make the code used in our analyses publicly available. These results offer practical guidance for efficient emulator design and simulation budgeting in future cosmological analyses.

Simulation budgeting for hybrid effective field theories

TL;DR

This work establishes a practical framework for budgeting N-body simulations to train hybrid effective field theory (HEFT) emulators for large-scale structure, balancing perturbation theory with N-body displacements to reach ~1% accuracy in the targeted k-range. By quantifying survey-driven statistical and modeling uncertainties (including intrinsic alignments and baryonic feedback) and testing higher-order bias terms, the authors derive explicit simulation requirements (box size, particle load, starting redshift) and present a tiered emulator training strategy across an 8-parameter cosmology space that includes CDM+. They demonstrate, via surrogate modeling and targeted use cases (DES Y6 GGL, SO CMB lensing, Roman cosmic shear, and steel samples), that emulators with sub-percent precision are achievable with on the order of a few hundred to a few thousand simulations, depending on the parameter volume and . The results provide concrete budgeting guidelines for efficient HEFT emulator design in future cosmological analyses. The methodology combines Lagrangian perturbation theory, HEFT, variance-reduction techniques (ZCV), and a PCA+PCE emulator to map cosmology-dependent power spectra across scales relevant for 2-point statistics.

Abstract

In this work, we forecast the number of, and requirements on, N-body simulations needed to train hybrid effective field theory (HEFT) emulators for a range of use cases, using a hybrid of HMcode and perturbation theory as a surrogate model. Our accuracy goals, determined with careful consideration of statistical and systematic uncertainties, are accurate in the high-likelihood range of cosmological parameters, and accurate over a broader parameter space volume for and . Focusing in part on the 8-parameter CDM+ cosmological model, we find that simulations are required to meet our error goals over our wide parameter space, including models with rapidly evolving dark energy, given our simulation and emulator recommendations. For a more restricted parameter space volume, as few as 80 simulations are sufficient. We additionally present simulation forecasts for example use cases, and make the code used in our analyses publicly available. These results offer practical guidance for efficient emulator design and simulation budgeting in future cosmological analyses.
Paper Structure (25 sections, 27 equations, 17 figures, 3 tables)

This paper contains 25 sections, 27 equations, 17 figures, 3 tables.

Figures (17)

  • Figure 1: Estimated theoretical error in the matter power spectrum due to neglected two-loop contributions. The dashed lines are $\propto k^4$, and the solid lines are given by $P_{\rm 2-loop}/P_{\rm tree} \approx \left[(P_{\rm 1-loop}-P_{\rm tree})/{P_{\rm tree}}\right]^2$, with one counter-term. Left: errors at four different redshifts for our fiducial cosmology. Right: errors for our fiducial cosmological parameters, but at noted values of $\sigma_8$.
  • Figure 2: Forecasted fractional uncertainties on angular galaxy and CMB lensing, and galaxy clustering, auto and cross-spectra for current and future surveys in our fiducial cosmology. For DESI, we use an extended LRG-like sample with a Gaussian $dN/dz$, with $\mu_z=0.633$, $\sigma_z=0.077$, large-scale bias $b=2$ and number density $\bar{n}_{\theta}=311\,{\rm deg}^{-2}$. We additionally assume $f_{\rm sky}=0.4$. For Rubin shear, we assume $\sigma_\epsilon = 0.26$, $n_{\rm eff}=5.54\,{\rm arcmin}^{-2}$, and $f_{\rm sky}=0.35$. For our Rubin lensing source sample, we use the second-highest-redshift LSST source bin Fang_2023. For SO, we use the publicly available lensing noise curve. We set $\Delta \ell = \sqrt{900+10\,\ell}$. A secondary $k$-axis, computed using $k = (\ell+0.5)/\chi(z=0.633)$, is provided for orientation.
  • Figure 3: Comparison between the counterterm model in eq. \ref{['eqn:ctrterm']}, the SP$(k)$ model, the $A_{\rm mod}$ model given by eq. \ref{['eqn:Amod']}, and measurements from two hydrodynamical simulations at $z=0.25$. The top panel shows the ratios of the power spectra with feedback to those ignoring feedback. The bottom panel shows the errors arising from ignoring feedback (solid), and from modeling feedback using the $A_{\rm mod}$ (dashed), the counterterm model (dot-dashed), and the SP$(k)$ model (dotted). The shaded grey band shows $\pm 1\%$. The blue and orange curves represent measurements from and fits to the BAHAMAS $T_{\rm heat}=7.6$ and C-OWLS with $T_{\rm heat}=8.7$ hydrodynamical simulations, respectively. The counterterm model fit to C-OWLS is plotted in a darker orange because it appears identical to the $A_{\rm mod}$ fit by eye. The measured power spectra were obtained from the https://powerlib.strw.leidenuniv.nl/van_Daalen_2019 and provide a crude estimate of the current range of uncertainty in the models. We expect feedback to be stronger at lower $z$ and weaker at higher $z$.
  • Figure 4: Plots showing HEFT (with quadratic bias) fits to spectra of mock galaxy catalogs with clustering similar to DESI LRGs in both $\Lambda$CDM and $w_0w_a\mathrm{CDM}+m_\nu$ cosmologies. The lower panels show the fractional errors with 1% and 3% bands shaded in grey. The left panels show the average of the ZCV variance-reduced galaxy auto- and cross-spectra measured from 13 of the base AbacusSummitMaksimova_2021Garrison_2021 simulations (which assume $\Lambda$CDM). The right panel shows the HEFT fits to mock galaxy auto- and cross-spectra measured from a single $w_0w_a{\rm CDM}+m_{\nu}$ simulation. These mock data are noticeably noisier than their $\Lambda$CDM counterparts due to the use of a single, smaller box. The best-fit biases for the Abacus simulations and $w_0w_a$ simulation are $b_{1}=1.4,\,b_2=2.6\times 10^{-2}, \,b_s=-0.27,\,b_{\nabla^2}=-2.8\times 10^{-2},\,{\rm sn}=73$ and $b_{1}=1.76,\,b_2=2.2, \,b_s=-1.44,\,b_{\nabla^2}=-0.45,\,{\rm sn}=212$ respectively.
  • Figure 5: Plots showing the gains from including successively higher order bias terms in HEFT fitted simultaneously to $P_{gg}(k)$ and $P_{gm}(k)$ for a single $\Lambda$CDM simulation at $z=0.53$. We choose $k_{\rm max}=0.6\,h\,\mathrm{Mpc}^{-1}$ for the quadratic and cubic models and $k_{\rm max}=0.15\,h\,\mathrm{Mpc}^{-1}$ for the linear model. Higher $k$ points are included in the fit, but are weighted such that $k\leq k_{\rm max}$ points are prioritized. The top left and right panels show the measured and fit halo auto- and cross-spectra, respectively, while the bottom panels showcase the fractional error in the HEFT fits. The impressive gain in accuracy in going from linear to quadratic bias is evident, particularly in the bottom panels. The bottom panels also showcase noticeable improvements in the halo auto- and cross-spectrum fits at the highest $k$ when including the cubic operator $:\!\!\delta^3\!\!:$. The light and dark grey shaded bands in the lower panels indicate $3\%$ and $1\%$ error, respectively.
  • ...and 12 more figures