Table of Contents
Fetching ...

Biology-informed neural networks learn nonlinear representations from omics data to improve genomic prediction and interpretability

Katiana Kontolati, Rini Jasmine Gladstone, Ian Davis, Ethan Pickering

TL;DR

Biology-informed neural networks (BINNs) address limitations of genotype-to-phenotype models by injecting biological inductive biases derived from multi-omics and pathway knowledge into neural architectures. The authors train BINNs on transcriptomics- and metabolomics-informed designs, enabling genotype-only inference at deployment while leveraging training-time omics to improve accuracy and uncover nonlinear relationships. They demonstrate up to 56% gains in test-set rank correlation and a 75% reduction in prediction error, and show latent variables that align with intermediate biology, supporting mechanistic interpretability. The work offers a practical framework for genomic selection in crops, guiding gene and pathway prioritization and enabling scalable, biology-aware design.

Abstract

We extend biologically-informed neural networks (BINNs) for genomic prediction (GP) and selection (GS) in crops by integrating thousands of single-nucleotide polymorphisms (SNPs) with multi-omics measurements and prior biological knowledge. Traditional genotype-to-phenotype (G2P) models depend heavily on direct mappings that achieve only modest accuracy, forcing breeders to conduct large, costly field trials to maintain or marginally improve genetic gain. Models that incorporate intermediate molecular phenotypes such as gene expression can achieve higher predictive fit, but they remain impractical for GS since such data are unavailable at deployment or design time. BINNs overcome this limitation by encoding pathway-level inductive biases and leveraging multi-omics data only during training, while using genotype data alone during inference. Applied to maize gene-expression and multi-environment field-trial data, BINN improves rank-correlation accuracy by up to 56% within and across subpopulations under sparse-data conditions and nonlinearly identifies genes that GWAS/TWAS fail to uncover. With complete domain knowledge for a synthetic metabolomics benchmark, BINN reduces prediction error by 75% relative to conventional neural nets and correctly identifies the most important nonlinear pathway. Importantly, both cases show highly sensitive BINN latent variables correlate with the experimental quantities they represent, despite not being trained on them. This suggests BINNs learn biologically-relevant representations, nonlinear or linear, from genotype to phenotype. Together, BINNs establish a framework that leverages intermediate domain information to improve genomic prediction accuracy and reveal nonlinear biological relationships that can guide genomic selection, candidate gene selection, pathway enrichment, and gene-editing prioritization.

Biology-informed neural networks learn nonlinear representations from omics data to improve genomic prediction and interpretability

TL;DR

Biology-informed neural networks (BINNs) address limitations of genotype-to-phenotype models by injecting biological inductive biases derived from multi-omics and pathway knowledge into neural architectures. The authors train BINNs on transcriptomics- and metabolomics-informed designs, enabling genotype-only inference at deployment while leveraging training-time omics to improve accuracy and uncover nonlinear relationships. They demonstrate up to 56% gains in test-set rank correlation and a 75% reduction in prediction error, and show latent variables that align with intermediate biology, supporting mechanistic interpretability. The work offers a practical framework for genomic selection in crops, guiding gene and pathway prioritization and enabling scalable, biology-aware design.

Abstract

We extend biologically-informed neural networks (BINNs) for genomic prediction (GP) and selection (GS) in crops by integrating thousands of single-nucleotide polymorphisms (SNPs) with multi-omics measurements and prior biological knowledge. Traditional genotype-to-phenotype (G2P) models depend heavily on direct mappings that achieve only modest accuracy, forcing breeders to conduct large, costly field trials to maintain or marginally improve genetic gain. Models that incorporate intermediate molecular phenotypes such as gene expression can achieve higher predictive fit, but they remain impractical for GS since such data are unavailable at deployment or design time. BINNs overcome this limitation by encoding pathway-level inductive biases and leveraging multi-omics data only during training, while using genotype data alone during inference. Applied to maize gene-expression and multi-environment field-trial data, BINN improves rank-correlation accuracy by up to 56% within and across subpopulations under sparse-data conditions and nonlinearly identifies genes that GWAS/TWAS fail to uncover. With complete domain knowledge for a synthetic metabolomics benchmark, BINN reduces prediction error by 75% relative to conventional neural nets and correctly identifies the most important nonlinear pathway. Importantly, both cases show highly sensitive BINN latent variables correlate with the experimental quantities they represent, despite not being trained on them. This suggests BINNs learn biologically-relevant representations, nonlinear or linear, from genotype to phenotype. Together, BINNs establish a framework that leverages intermediate domain information to improve genomic prediction accuracy and reveal nonlinear biological relationships that can guide genomic selection, candidate gene selection, pathway enrichment, and gene-editing prioritization.
Paper Structure (21 sections, 11 equations, 3 figures, 3 tables, 1 algorithm)

This paper contains 21 sections, 11 equations, 3 figures, 3 tables, 1 algorithm.

Figures (3)

  • Figure 1: Biology-informed neural network framework embed domain knowledge for enhanced genomic prediction and learning nonlinear biological relationships.$a)$ Conventional G2P models use genotype alone, leaving rich functional knowledge underutilized. BINNs embed curated biology (from RNA-seq (expression), methylomics (DNA methylation), metabolomics, KEGG pathway annotations, and proteomics) directly into the architecture in the form of pathway structure, regulatory priors, and sparsity constraints, to boost predictive accuracy while preserving practical utility. $b)$ Four representative cases where GWAS, TWAS, and BINN are valid for analyzing genomic, transcriptomic, and phenomic datasets. Only BINN permits association under general nonlinearity (with the assumption the trained BINN model is accurate).
  • Figure 2: BINNs improve prediction accuracy and interpretability through utiliziation of gene expression in genotype to phenotype modeling. (a) Schematic of the transcriptomics-derived BINN architecture. The SNP marker and gene expression data are utilized for feature selection, which serves to sparsify the connections between the input and intermediate layers of the BINN architecture. The marker data for each gene is passed through pathway subnetworks at the intermediate layer and the outputs from this layer is passed through a non-linear integrator network to predict the phenotype values. The G2P and B2P models are both linear models. (b) Predicted vs. observed phenotype days to anthesis in MI across four subpopulations (SS, NSS, IDT, and Others). Points show G2P (GBLUP on genotype; pink triangles) and G2B2P (BINN with expression-informed sparsity; blue circles). An ordinary least-squares (OLS) fit is overlaid and the Spearman correlation ($\rho$) is reported in the legend. (c) Test Spearman correlation distributions for Silking NE, comparing three predictive models (G2P, G2B2P and B2P: Ridge regression on expression, green) across five independent 20/80% train/test splits, each with five‑fold cross‑validation to capture both split‑level and CV‑level variability. Each subplot shows results for all lines pooled (“all”) and for each of the seven distinct subpopulations (SS, NSS, IDT, Popcorn, Sweet corn, Tropical and Others). The pink dashed line marks the median of the G2P baseline; the dashed line for G2B2P is colored green when its median exceeds that baseline or red when it falls below. Legends report the median percent change of G2B2P relative to G2P and titles show the number of test samples. (d) Test Spearman correlation distributions for leave-one-population-out experiments for Silking NE, comparing the same three predictive models. Each subplot shows results for the predictions on six distinct subpopulations using the models that are trained by leaving out the corresponding subpopulation from the training data. (e) Predicted latent variables vs. observed gene expression for four representative high-correlation genes per phenotype. (f) Aggregated absolute phenotypic change for 30 representative genes - 15 with high absolute Pearson correlation and 15 close to zero. (g) Aggregated absolute phenotypic change per perturbed gene, with a threshold capturing the BINN-derived 100 most significant genes which includes zap1, zmm15, zcn14 and zcn8.
  • Figure 3: BINNs significantly outperform baselines in the sparse data limit $n<p$. (a) Schematic of the BINN architecture for the shoot‐branching network. Gene inputs are routed through four biologically‐annotated pathway subnetworks corresponding to auxin (A), sucrose (S), cytokinins (CK), and strigolactones (SL), then combined in a final integrator to predict bud‐outgrowth time. (b) Test‐set performance (MSE) for six models: ridge regression (RR), a generic fully‐connected network (FCN), BINN trained with standard MSE (BINN‐MSE), and BINN trained with the biologically‐informed soft‐constraint loss (BINN‐soft) with a varying fraction of known intermediate trait measurements (100%, 50% and 10%), evaluated across nine training‐set sizes logarithmically spaced between 500 and 20,000 samples. Boxplots display the distribution across five random initializations. The vertical black dashed line at $n=1,600$ denotes the transition from sparse ($n<p$) to plentiful ($n>p$) data regimes. (c) Predicted vs observed phenotype across four training sizes ($n$ = 500, 1,994, 7,953, 20,000) for three models: RR (purple circles), FCN (black triangles), and BINN (red squares). Each panel shows the identity line (y = x), and in-panel legends report Pearson correlation (r) for each model. (d) Scatter plots of predicted versus true latent values for each intermediate trait (A, S, CK, SL) under the BINN MSE (blue) and the BINN soft model (red). The fitted regression line and Pearson correlation coefficient, r, are overlaid. (e) Aggregated absolute phenotypic change per perturbed intermediate trait.