Fitness inference tested by in silico population genetics
Hong-Li Zeng, Yu-Han Huang, John Barton, Erik Aurell
TL;DR
The paper tackles whether fitness parameters and genotype fitness order can be inferred from time-series, whole-genome data under selection, mutation, and recombination. It compares two complementary inference schemes, marginal path likelihood ($MPL$) and transient quasi-linkage equilibrium ($tQLE$), applying them to simulated populations with additive and pairwise epistatic fitness. The authors map parameter regimes where fitness inference is feasible and examine recovery of additive and epistatic components, finding that MPL and $tQLE$ largely agree across broad ranges, especially for ranking top fitness sequences. This work provides a practical framework for planning real-data analyses in pathogens and ancient-DNA contexts, clarifying when fitness inference is possible and highlighting the usefulness of focusing on genotype ranks.
Abstract
We consider populations evolving according to natural selection, mutation, and recombination, and assume that the genomes of all or a representative selection of individuals are known. We pose the problem if it is possible to infer fitness parameters and genotype fitness order from such data. We tested this hypothesis in simulated populations. We delineate parameter ranges where this is possible and other ranges where it is not.Our work provides a framework for determining when fitness inference is feasible from population-wide, whole-genome, time-stratified data and highlights settings where it is not. We give a brief survey of biological model organisms and human pathogens that fit into this framework.
