Table of Contents
Fetching ...

LyαNNA II: Field-level inference with noisy Lyα forest spectra

Parth Nayak, Michael Walther, Daniel Gruen

TL;DR

This work advances field-level inference for the Lyα forest by incorporating realistic nuisances such as noise and finite spectral resolution into the LyαNNA framework. It trains nSansa, a 1D ResNet, to compress spectra into a 2D summary targeting the IGM's power-law temperature-density relation parameters $(T_0,\gamma)$, and compares Gaussian-likelihood with likelihood-free (DELFI) inference. Across a broad range of continuum-to-noise ratios, the NN-based approach yields sizable gains in posterior precision over traditional summaries (TPS+TPDF), with FoM improvements from roughly $1.65\times$ to $2.12\times$. The Gaussian posterior is found to be a good approximation even for DELFI in this setting, highlighting the practical value of field-level, machine-learned summaries for noisy spectra and guiding future work to include additional observational systematics.

Abstract

Deep learning (DL) has been shown to outperform traditional, human-defined summary statistics of the Lyα forest in constraining key astrophysical and cosmological parameters owing to its ability to tap into the realm of non-Gaussian information. An understanding of the impact of nuisance effects such as noise on such field-level frameworks, however, still remains elusive. In this work we conduct a systematic investigation into the efficacy of DL inference from noisy Lyα forest spectra. Building upon our previous, proof-of-concept framework (Nayak et al. 2024) for pure spectra, we constructed and trained a ResNet neural network using labeled mock data from hydrodynamical simulations with a range of noise levels to optimally compress noisy spectra into a novel summary statistic that is exclusively sensitive to the power-law temperature-density relation of the intergalactic medium. We fit a Gaussian mixture surrogate with 23 components through our labels and summaries to estimate the joint data-parameter distribution for likelihood free inference, in addition to performing inference with a Gaussian likelihood. The posterior contours in the two cases agree well with each other. We compared the precision and accuracy of our posterior constraints with a combination of two human defined summaries (the 1D power spectrum and PDF of the Lyα transmission) that have been corrected for noise, over a wide range of continuum-to-noise ratios (CNR) in the likelihood case. We found a gain in precision in terms of posterior contour area with our pipeline over the said combination of 65% (at a CNR of 20 per 6 km/s) to 112% (at 200 per 6 km/s). While the improvement in posterior precision is not as large as in the noiseless case, these results indicate that DL still remains a powerful tool for inference even with noisy, real-world datasets.

LyαNNA II: Field-level inference with noisy Lyα forest spectra

TL;DR

This work advances field-level inference for the Lyα forest by incorporating realistic nuisances such as noise and finite spectral resolution into the LyαNNA framework. It trains nSansa, a 1D ResNet, to compress spectra into a 2D summary targeting the IGM's power-law temperature-density relation parameters , and compares Gaussian-likelihood with likelihood-free (DELFI) inference. Across a broad range of continuum-to-noise ratios, the NN-based approach yields sizable gains in posterior precision over traditional summaries (TPS+TPDF), with FoM improvements from roughly to . The Gaussian posterior is found to be a good approximation even for DELFI in this setting, highlighting the practical value of field-level, machine-learned summaries for noisy spectra and guiding future work to include additional observational systematics.

Abstract

Deep learning (DL) has been shown to outperform traditional, human-defined summary statistics of the Lyα forest in constraining key astrophysical and cosmological parameters owing to its ability to tap into the realm of non-Gaussian information. An understanding of the impact of nuisance effects such as noise on such field-level frameworks, however, still remains elusive. In this work we conduct a systematic investigation into the efficacy of DL inference from noisy Lyα forest spectra. Building upon our previous, proof-of-concept framework (Nayak et al. 2024) for pure spectra, we constructed and trained a ResNet neural network using labeled mock data from hydrodynamical simulations with a range of noise levels to optimally compress noisy spectra into a novel summary statistic that is exclusively sensitive to the power-law temperature-density relation of the intergalactic medium. We fit a Gaussian mixture surrogate with 23 components through our labels and summaries to estimate the joint data-parameter distribution for likelihood free inference, in addition to performing inference with a Gaussian likelihood. The posterior contours in the two cases agree well with each other. We compared the precision and accuracy of our posterior constraints with a combination of two human defined summaries (the 1D power spectrum and PDF of the Lyα transmission) that have been corrected for noise, over a wide range of continuum-to-noise ratios (CNR) in the likelihood case. We found a gain in precision in terms of posterior contour area with our pipeline over the said combination of 65% (at a CNR of 20 per 6 km/s) to 112% (at 200 per 6 km/s). While the improvement in posterior precision is not as large as in the noiseless case, these results indicate that DL still remains a powerful tool for inference even with noisy, real-world datasets.
Paper Structure (16 sections, 7 equations, 11 figures, 1 table)

This paper contains 16 sections, 7 equations, 11 figures, 1 table.

Figures (11)

  • Figure 1: The sample of training, validation, and test labels in our mock dataset along with the fiducial TDR model in the $(T_0,\gamma)$ as well as the rescaled $(\tilde{T}_0, \tilde{\gamma})$ space. The gray shaded region indicates our prior for the likelihood analysis as well as for the density estimation likelihood free inference. The exact sampling strategy is described in Appendix \ref{['app:sampling']}.
  • Figure 2: Traditional summary statistics (TPS on the left, TPDF on the right) estimated for the fiducial thermal state of our simulation box and with the mean transmission fixed to its observed value. The raw estimators correspond to a noise level of CNR$_6=30$ ($\sigma_\mathrm{p}=0.014$). The corresponding pure statistics are computed from 100,000 noiseless spectra (with infinite spectral resolution in TPS, with a resolution $R_\mathrm{FWHM}=6000$ for TPDF). The errors correspond to $N_\mathrm{s}=100$ spectra. The gray regions in TPDF correspond to our cuts due to edge effects in the deconvoled estimator.
  • Figure 3: The full correlation matrices of the concatenated summary vector for four different noise levels. The TPS is the corrected vector $\hat{\mathbb{P}}_\mathrm{pure}$ with 122 $k$-modes and the TPDF is the deconvolved, cropped $\hat{p}_\mathrm{pure}$ with 45 bins. For each of the four matrices, the two blocks on the principal diagonal are the individual correlation matrices of TPS (top left) and TPDF (bottom right) and the other two blocks show the cross correlation of TPS and TPDF. The different blocks have been resized differently to have an equal area on the plot.
  • Figure 4: Architecture of nSansa. An input spectrum of size 256 pixels is fed into the network that contains a total of 6 residual blocks and extracts useful features from the field. Batch normalization and dropout are used for regularization and average pooling is used for downsampling. The output of the residual part is flattened, concatenated with the $\sigma_\mathrm{p}$ query, and then fed into a hidden nonlinear layer with $N_\mathrm{d}$ units. Finally, a linear layer with 5 nodes (2 for the summary vector, 3 for its covariance) acts as the output layer.
  • Figure 5: An example of typical learning curves of the nSansa architecture. The gray band in the $\chi^2$ panel indicates our tolerance of $\epsilon=0.05$. Here the minimum of the validation loss is reached at epoch $j^*=540$ and the $\chi^2$ is simultaneously within our tolerance.
  • ...and 6 more figures