Table of Contents
Fetching ...

Fitness inference tested by in silico population genetics

Hong-Li Zeng, Yu-Han Huang, John Barton, Erik Aurell

TL;DR

The paper tackles whether fitness parameters and genotype fitness order can be inferred from time-series, whole-genome data under selection, mutation, and recombination. It compares two complementary inference schemes, marginal path likelihood ($MPL$) and transient quasi-linkage equilibrium ($tQLE$), applying them to simulated populations with additive and pairwise epistatic fitness. The authors map parameter regimes where fitness inference is feasible and examine recovery of additive and epistatic components, finding that MPL and $tQLE$ largely agree across broad ranges, especially for ranking top fitness sequences. This work provides a practical framework for planning real-data analyses in pathogens and ancient-DNA contexts, clarifying when fitness inference is possible and highlighting the usefulness of focusing on genotype ranks.

Abstract

We consider populations evolving according to natural selection, mutation, and recombination, and assume that the genomes of all or a representative selection of individuals are known. We pose the problem if it is possible to infer fitness parameters and genotype fitness order from such data. We tested this hypothesis in simulated populations. We delineate parameter ranges where this is possible and other ranges where it is not.Our work provides a framework for determining when fitness inference is feasible from population-wide, whole-genome, time-stratified data and highlights settings where it is not. We give a brief survey of biological model organisms and human pathogens that fit into this framework.

Fitness inference tested by in silico population genetics

TL;DR

The paper tackles whether fitness parameters and genotype fitness order can be inferred from time-series, whole-genome data under selection, mutation, and recombination. It compares two complementary inference schemes, marginal path likelihood () and transient quasi-linkage equilibrium (), applying them to simulated populations with additive and pairwise epistatic fitness. The authors map parameter regimes where fitness inference is feasible and examine recovery of additive and epistatic components, finding that MPL and largely agree across broad ranges, especially for ranking top fitness sequences. This work provides a practical framework for planning real-data analyses in pathogens and ancient-DNA contexts, clarifying when fitness inference is possible and highlighting the usefulness of focusing on genotype ranks.

Abstract

We consider populations evolving according to natural selection, mutation, and recombination, and assume that the genomes of all or a representative selection of individuals are known. We pose the problem if it is possible to infer fitness parameters and genotype fitness order from such data. We tested this hypothesis in simulated populations. We delineate parameter ranges where this is possible and other ranges where it is not.Our work provides a framework for determining when fitness inference is feasible from population-wide, whole-genome, time-stratified data and highlights settings where it is not. We give a brief survey of biological model organisms and human pathogens that fit into this framework.
Paper Structure (12 sections, 12 equations, 9 figures, 2 tables)

This paper contains 12 sections, 12 equations, 9 figures, 2 tables.

Figures (9)

  • Figure 1: Scatter plots for the reconstructed additive fitness $f_i^*$ versus the ground truth $f_i$ by the tQLE (blue) and MPL (red) approaches, displaying the effects of increasing mutation rate (parameter $\mu$) and increasing fluctuating additive fitness (parameter $\sigma(f_i)$). Upper: $\mu = 0.003$, $\sigma(f_i) = 0.01$, middle: $\mu = 0.01$, $\sigma(f_i) = 0.01$, bottom: $\mu = 0.01$, $\sigma(f_i) = 0.05$. Fitness parameters $f_i$ are drawn randomly from a Gaussian distribution with mean zero and standard deviation $\sigma(f_i)$. In these simulations the underlying ground truth is only additive fitness (no epistatic fitness), and corresponding model parameter therefore $\sigma(f_{ij})=0$. Other parameter values as indicated in Table \ref{['tab:parameter-values']}.
  • Figure 2: Scatter plots for total fitness of sequences inferred by MPL and tQLE against ground-truth fitness values displaying the effects of increasing mutation rate (parameter $\mu$) and increasing fluctuating additive fitness (parameter $\sigma(f_i)$, with the same parameter choices as in Fig. \ref{['fig:scatter_plts_for_fi']}. Red triangles represent MPL results, and blue stars represent tQLE results. In the tests reported here only the additive fitness part of tQLE is taken into account i.e. inferred epistatic fitness $f_{ij}^*$ is set to zero throughout, including as it enters in the formula for inferred additive fitness \ref{['eq:f-1-inference-QLE']}. In these simulations the underlying ground truth is only additive fitness (no epistatic fitness), and corresponding model parameter therefore $\sigma(f_{ij})=0$. Other parameter values as indicated in Table \ref{['tab:parameter-values']}.
  • Figure 3: Scatter plot of fitness rank of top 5% most fit sequences as inferred using MPL and tQLE against ground truth. As in Fig. \ref{['fig:scatter_plts_for_estimated_additive_fitness']} ground truth fitness is due to additive fitness alone, and tQLE is used without taking epistatic fitness into account. Each panel displays the inferred rank (y-axis) of the top 5% ground-truth fitness sequences (x-axis) for MPL (red triangles) and tQLE (blue stars). Mutation ($\mu$) and additive fitness ($\sigma(f_i)$) parameters the same as in Figs \ref{['fig:scatter_plts_for_fi']} and \ref{['fig:scatter_plts_for_estimated_additive_fitness']}, other parameter values as indicated in Table \ref{['tab:parameter-values']}. As additive effects strengthen, both methods achieve improved rank accuracy. MPL slightly outperforms tQLE under stronger selection, while the latter exhibits greater stability under lower mutations.
  • Figure 4:
  • Figure 5: Performance comparison between tQLE and MPL in fitness inference under different $\sigma(f_i)$. Each panel shows the Spearman correlation between the ranks of inferred fitness values compared to the ground truth for 30 independent realizations, with increasing additive fitness variability (parameter $\sigma(f_i)$). Parameters are the same as in Fig. \ref{['fig:histograms_for_differnt_fis']} and as indicated in Table \ref{['tab:parameter-values']}. Dashed lines represent the correlations of ranks of the top 5% ground-truth fittest sequences inferred by MPL (red) and tQLE (blue), while solid lines correspond to the rank of all sequences. Performance of tQLE increases with increasing number of replicates while MPL is relatively insensitive to number of replicates. For lowest additive fitness (top panel), tQLE is better at inferring fitness across all sequences beyond about 5 replicates, and for the top 5% of sequences at any number of replicates. At intermediate additive fitness (middle panel), MPL outperforms tQLE for all sequences while tQLE is marginally better for the top 5%, at sufficiently large number of replicates. For high additive fitness (bottom panel) MPL works consistently better.
  • ...and 4 more figures