Table of Contents
Fetching ...

Black Hole-Host Galaxy Correlations with Machine Learning: A Comparative Study of Illustris, TNG, and EAGLE

Jacob Reinheimer, Yuan Li, Trung Ha, Melanie Habouzit, Brandon M. Matthews, George Blaney

TL;DR

This study examines SMBH–host galaxy scaling relations at $z=0$ across Illustris, TNG, and EAGLE using multiple ML regressors to quantify predictive power and non-linearities. It shows that ML methods, particularly Multi-layer Perceptron networks, generally outperform linear fits, with the strongest single-relator link being $M_{ m BH}$–$\sigma$ in all simulations and $M_{ m BH}$–$M_{\star}$ strongest in TNG. The results reveal substantial simulation-dependent differences in the strength and shape of SMBH–host correlations, likely driven by sub-grid feedback models, with EAGLE showing the weakest couplings. The work highlights the value of a multi-dimensional, ML-based framework to quantify SMBH–host coevolution and to compare simulations with observations as data sets grow.

Abstract

Supermassive black holes (SMBHs) are known to correlate with many properties of their host galaxies, but we do not fully understand these correlations. The strengths (tightness) of these correlations have also been widely debated. In this work, we explore SMBH-host relations in three state-of-the-art cosmological simulations: Illustris, TNG, and EAGLE. Using a variety of machine learning regressors, we measure the scaling relations between black hole mass ($M_{\rm BH}$) and galaxy properties including stellar velocity dispersion ($σ$), stellar mass ($M_{\star}$), dark matter halo mass ($M_{\rm Halo}$), and the Sersic index. We find that machine learning regressors provide predictive capabilities superior to linear regression in many scaling relations in simulations, and Multi-layer Perceptron (MLP) regressor has the strongest performance. SMBH-host relations have different strengths in different simulations as a result of their sub-grid models. Similar to the observations, the $M_{\rm BH} $-$σ$ relation is a strong correlation in all simulations, but in TNG, the $M_{\rm BH} $-$M_{\star}$ relation is even tighter than $M_{\rm BH} $-$σ$. EAGLE produces the weakest SMBH-host correlations among all simulations. Low mass SMBHs tend to be poorly correlated with their host galaxies, but including them can still help machines better grasp the correlations in Illustris and TNG. Combining galaxy properties that strongly correlate with $M_{\rm BH} $ but poorly correlate with each other can improve MLP's performance. $M_{\rm BH} $ is most accurately predicted when all galaxy properties are included in the training, suggesting that SMBH-host correlations are fundamentally multi-dimensional in these simulations.

Black Hole-Host Galaxy Correlations with Machine Learning: A Comparative Study of Illustris, TNG, and EAGLE

TL;DR

This study examines SMBH–host galaxy scaling relations at across Illustris, TNG, and EAGLE using multiple ML regressors to quantify predictive power and non-linearities. It shows that ML methods, particularly Multi-layer Perceptron networks, generally outperform linear fits, with the strongest single-relator link being in all simulations and strongest in TNG. The results reveal substantial simulation-dependent differences in the strength and shape of SMBH–host correlations, likely driven by sub-grid feedback models, with EAGLE showing the weakest couplings. The work highlights the value of a multi-dimensional, ML-based framework to quantify SMBH–host coevolution and to compare simulations with observations as data sets grow.

Abstract

Supermassive black holes (SMBHs) are known to correlate with many properties of their host galaxies, but we do not fully understand these correlations. The strengths (tightness) of these correlations have also been widely debated. In this work, we explore SMBH-host relations in three state-of-the-art cosmological simulations: Illustris, TNG, and EAGLE. Using a variety of machine learning regressors, we measure the scaling relations between black hole mass () and galaxy properties including stellar velocity dispersion (), stellar mass (), dark matter halo mass (), and the Sersic index. We find that machine learning regressors provide predictive capabilities superior to linear regression in many scaling relations in simulations, and Multi-layer Perceptron (MLP) regressor has the strongest performance. SMBH-host relations have different strengths in different simulations as a result of their sub-grid models. Similar to the observations, the - relation is a strong correlation in all simulations, but in TNG, the - relation is even tighter than -. EAGLE produces the weakest SMBH-host correlations among all simulations. Low mass SMBHs tend to be poorly correlated with their host galaxies, but including them can still help machines better grasp the correlations in Illustris and TNG. Combining galaxy properties that strongly correlate with but poorly correlate with each other can improve MLP's performance. is most accurately predicted when all galaxy properties are included in the training, suggesting that SMBH-host correlations are fundamentally multi-dimensional in these simulations.
Paper Structure (23 sections, 6 figures, 2 tables)

This paper contains 23 sections, 6 figures, 2 tables.

Figures (6)

  • Figure 1: $M_{\rm BH}$ distributions at $z=0$ for the Illustris (Blue), TNG (Orange), and EAGLE (Green) simulations. Each vertical line represents the median value of the $M_{\rm BH}$ in that simulation sample (the most massive 3607 galaxies), and is used to divide each sample into high- and low-mass SMBHs for the analyses in Section \ref{['sec:mass_dependence']}.
  • Figure 2: SMBH--host galaxy correlations at $z=0$ in the Illustris (left), TNG (middle), and EAGLE (right) simulations. From top to bottom, we show the $M_{\rm BH} \text{--} \sigma$ the $M_{\rm BH}$--$M_{\star}$, the $M_{\rm BH}$--$M_{\rm Halo}$, and the $M_{\rm BH}$--$\acute{n}$ relations. Only the most massive 3607 galaxies in each simulation are included in our study, which corresponds to $M_{\star} \gtrsim 10^{10} M_{\odot}$ for EAGLE. The black solid line is a linear least squares regression fit, and the error bar in the bottom right describes one standard deviation errors.
  • Figure 3: The strength of SMBH-host correlations in Illustris (left), TNG (middle), and EAGLE (right) measured by machine learning regressors' ability to predict $M_{\rm BH}$. Small MSEs (errors in predicted $M_{\rm BH}$) correspond to tight correlations. We repeat the analysis 50 times for each regressor to quantify the uncertainty from the random training/testing split. Solid lines in the top panels connect the average MSEs of the 50 realizations while the shaded regions show the standard deviation. For clarity, we only show linear regressor and MLP in the top panels. In the bottom panels, the other regressors's performances are divided by MLP to show their relative strengths.
  • Figure 4: The tightness of SMBH-host relations measured with optimized MLP regressor for low- and high-mass SMBHs in Illustris (left), TNG (middle), and EAGLE (right). Low- and high-mass SMBHs are defined as SMBHs below and above the median $M_{\rm BH}$ in each sample (see Figure \ref{['fig:hist']}). Similar to Figure \ref{['fig:MSE_results']}, the lines connect the average MSEs of 50 iterations and the shaded regions represent the standard deviation. For comparison, we also show the full sample (same as the black line in Figure \ref{['fig:MSE_results']}) and a randomly selected half sample (see Section \ref{['sec:mass_dependence']} for details).
  • Figure 5: MLP regressor's measurement of SMBH-host relations when two galaxy properties are combined. The diagonal cells use only one galaxy property (same as the black line in Figure \ref{['fig:MSE_results']}).
  • ...and 1 more figures