Table of Contents
Fetching ...

Living on the edge: Testing for compact population features at the edges of parameter space

Asad Hussain, Maximiliano Isi, Aaron Zimmerman

TL;DR

This work tackles the bias and variance challenges that arise when inferring astrophysical populations with parameters restricted to bounded domains, such as black hole spins $\chi\in[0,1]$. It introduces a truncated Gaussian mixture model (TGMM) that partitions parameters into an Analytic Sector (with closed-form integrals against a truncated Gaussian kernel) and a Sampled Sector (handled via Monte Carlo), enabling boundary-faithful, efficient hierarchical inference. The method is validated with toy 1D examples and applied to gravitational-wave populations in the GWTC-3 catalog, reproducing known results (e.g., zero-spin fractions) while avoiding expensive reanalysis. The approach reduces boundary-induced biases and variance, can scale to large catalogs, and is implemented in open-source packages gravpop and truncatedgaussianmixtures for practical use in GW population studies.

Abstract

Many astrophysical population studies involve parameters that exist on a bounded domain, such as the dimensionless spins of black holes or the eccentricities of planetary orbits, both of which are confined to $[0, 1]$. In such scenarios, we often wish to test for distributions clustered near a boundary, e.g., vanishing spin or orbital eccentricity. Conventional approaches -- whether based on Monte Carlo, kernel density estimators, or machine-learning techniques -- often suffer biases at the boundaries. These biases stem from sparse sampling near the edge, kernel-related smoothing, or artifacts introduced by domain transformations. We introduce a truncated Gaussian mixture model framework that substantially mitigates these issues, enabling accurate inference of narrow, edge-dominated population features. While our method has broad applications to many astronomical domains, we consider gravitational wave catalogs as a concrete example to demonstrate its power. In particular, we maintain agreement with published constraints on the fraction of zero-spin binary black hole systems in the GWTC-3 catalog -- results originally derived at much higher computational cost through dedicated reanalysis of individual events in the catalog. Our method can achieve similarly reliable results with a much lower computational cost. The method is publicly available in the open-source packages gravpop and truncatedgaussianmixtures.

Living on the edge: Testing for compact population features at the edges of parameter space

TL;DR

This work tackles the bias and variance challenges that arise when inferring astrophysical populations with parameters restricted to bounded domains, such as black hole spins . It introduces a truncated Gaussian mixture model (TGMM) that partitions parameters into an Analytic Sector (with closed-form integrals against a truncated Gaussian kernel) and a Sampled Sector (handled via Monte Carlo), enabling boundary-faithful, efficient hierarchical inference. The method is validated with toy 1D examples and applied to gravitational-wave populations in the GWTC-3 catalog, reproducing known results (e.g., zero-spin fractions) while avoiding expensive reanalysis. The approach reduces boundary-induced biases and variance, can scale to large catalogs, and is implemented in open-source packages gravpop and truncatedgaussianmixtures for practical use in GW population studies.

Abstract

Many astrophysical population studies involve parameters that exist on a bounded domain, such as the dimensionless spins of black holes or the eccentricities of planetary orbits, both of which are confined to . In such scenarios, we often wish to test for distributions clustered near a boundary, e.g., vanishing spin or orbital eccentricity. Conventional approaches -- whether based on Monte Carlo, kernel density estimators, or machine-learning techniques -- often suffer biases at the boundaries. These biases stem from sparse sampling near the edge, kernel-related smoothing, or artifacts introduced by domain transformations. We introduce a truncated Gaussian mixture model framework that substantially mitigates these issues, enabling accurate inference of narrow, edge-dominated population features. While our method has broad applications to many astronomical domains, we consider gravitational wave catalogs as a concrete example to demonstrate its power. In particular, we maintain agreement with published constraints on the fraction of zero-spin binary black hole systems in the GWTC-3 catalog -- results originally derived at much higher computational cost through dedicated reanalysis of individual events in the catalog. Our method can achieve similarly reliable results with a much lower computational cost. The method is publicly available in the open-source packages gravpop and truncatedgaussianmixtures.
Paper Structure (24 sections, 73 equations, 11 figures, 1 table, 1 algorithm)

This paper contains 24 sections, 73 equations, 11 figures, 1 table, 1 algorithm.

Figures (11)

  • Figure 1: Bias and variance comparisons between different estimation methods in a toy example. Both panels show the same comparison: estimating the likelihood for a posterior with shape $\phi_{[0,1]}(0.4, 0.2)$ and a population model that gets increasingly compact near the edge $\phi_{[0,1]}(0, \sigma), \sigma \to 0$. The variance of the likelihood estimate (95% confidence interval shown here) can blow up for values of the population deviation parameter smaller than $\sigma \approx 0.025$ for 1000 samples. The confidence intervals of the (using bootstrapped estimates) are much tighter and around the correct value. Additionally techniques can fix this blow-up in the variance estimate, but can give biased estimates as we approach the edge. This efficiency and bias only gets worse with an increase in the number of dimensions.
  • Figure 2: Corner plot comparing the fit and the posterior samples for GW150914.
  • Figure 3: Comparison of integral estimates from both the method and when evaluated at the median hyper-parameters ($\boldsymbol{\Lambda}_0$) of our population model, but with $\mu_\chi = 0$ and $\sigma_\chi$ allowed to vary. The $2\sigma$ variance of the estimate is shown as blue bands. In both cases the estimate remains consistent with the estimate as the width of the spin magnitude population narrows. Left: Comparison of the marginalized likelihoods for GW150914. Right: Comparison of the detection efficiency.
  • Figure 4: The fractional bias between the between and integrals, relative to the result. Each point is the integral estimate evaluated at a hyper-posterior sample $\boldsymbol{\Lambda}_i$ from our full population inference using the model described in Table \ref{['tab:WidePriorModels']}. The error bars encapsulate the $2 \sigma$ deviation of the estimate as expected by the variance of the estimator. The error bars denote the $2 \sigma$ deviation of the estimate as expected by the variance of the estimator. We color those samples outside the $2 \sigma$ range green, and those which have $N_{\text{eff}}$ too small red. Left: The fractional bias in the marginalized likelihood estimate for GW150914. We find that only $\approx 2.1\%$ of samples are outside the $2 \sigma$ range. Right: The fractional bias in the detection efficiency. We find that only $\approx 0.3\%$ of samples are outside this $2 \sigma$ range.
  • Figure 5: For each hyper-posterior sample we compute the number of effective samples in the estimate for each event. Since our likelihood estimate is only valid if we have good convergence for all events we plot the smallest $N_{\text{eff}}$ for each sample on the $y$-axis (this is usually GW190517). We can see that for $\sigma_\chi < 0.15$ we are in a regime where the estimate has too few points. As such we only use posterior samples with $\sigma_\chi > 0.15$ to directly compare the and results.
  • ...and 6 more figures