Table of Contents
Fetching ...

Transforming Multi-Omics Integration with GANs: Applications in Alzheimer's and Cancer

Md Selim Reza, Sabrin Afroz, Mostafizer Rahman, Md Ashad Alam

TL;DR

This work tackles the challenges of integrating heterogeneous multi-omics data under limited samples and high noise by introducing Omics-GAN, a three-WGAN framework that generates high-quality synthetic profiles for mRNA, DNA methylation, and miRNA while preserving inter-omics relationships. Applied to ROSMAP (Alzheimer’s) and TCGA liver/colon cancer, Omics-GAN consistently improved disease outcome predictions (e.g., higher AUC in SVM classifiers) and revealed biologically meaningful features that persisted across original and synthetic data, validated by GO and KEGG enrichment analyses. The study further demonstrates the value of incorporating interaction networks, showing that true biological networks enhance predictive gains over random networks, and leverages molecular docking to propose drug repurposing candidates (Nilotinib for AD, Atovaquone for liver cancer, Tecovirimat for colon cancer). Collectively, Omics-GAN advances accurate biomarker discovery and therapeutic prioritization by generating faithful synthetic multi-omics data, offering a scalable path for precision medicine.

Abstract

Multi-omics data integration is crucial for understanding complex diseases, yet limited sample sizes, noise, and heterogeneity often reduce predictive power. To address these challenges, we introduce Omics-GAN, a Generative Adversarial Network (GAN)-based framework designed to generate high-quality synthetic multi-omics profiles while preserving biological relationships. We evaluated Omics-GAN on three omics types (mRNA, miRNA, and DNA methylation) using the ROSMAP cohort for Alzheimer's disease (AD) and TCGA datasets for colon and liver cancer. A support vector machine (SVM) classifier with repeated 5-fold cross-validation demonstrated that synthetic datasets consistently improved prediction accuracy compared to original omics profiles. The AUC of SVM for mRNA improved from 0.72 to 0.74 in AD, and from 0.68 to 0.72 in liver cancer. Synthetic miRNA enhanced classification in colon cancer from 0.59 to 0.69, while synthetic methylation data improved performance in liver cancer from 0.64 to 0.71. Boxplot analyses confirmed that synthetic data preserved statistical distributions while reducing noise and outliers. Feature selection identified significant genes overlapping with original datasets and revealed additional candidates validated by GO and KEGG enrichment analyses. Finally, molecular docking highlighted potential drug repurposing candidates, including Nilotinib for AD, Atovaquone for liver cancer, and Tecovirimat for colon cancer. Omics-GAN enhances disease prediction, preserves biological fidelity, and accelerates biomarker and drug discovery, offering a scalable strategy for precision medicine applications.

Transforming Multi-Omics Integration with GANs: Applications in Alzheimer's and Cancer

TL;DR

This work tackles the challenges of integrating heterogeneous multi-omics data under limited samples and high noise by introducing Omics-GAN, a three-WGAN framework that generates high-quality synthetic profiles for mRNA, DNA methylation, and miRNA while preserving inter-omics relationships. Applied to ROSMAP (Alzheimer’s) and TCGA liver/colon cancer, Omics-GAN consistently improved disease outcome predictions (e.g., higher AUC in SVM classifiers) and revealed biologically meaningful features that persisted across original and synthetic data, validated by GO and KEGG enrichment analyses. The study further demonstrates the value of incorporating interaction networks, showing that true biological networks enhance predictive gains over random networks, and leverages molecular docking to propose drug repurposing candidates (Nilotinib for AD, Atovaquone for liver cancer, Tecovirimat for colon cancer). Collectively, Omics-GAN advances accurate biomarker discovery and therapeutic prioritization by generating faithful synthetic multi-omics data, offering a scalable path for precision medicine.

Abstract

Multi-omics data integration is crucial for understanding complex diseases, yet limited sample sizes, noise, and heterogeneity often reduce predictive power. To address these challenges, we introduce Omics-GAN, a Generative Adversarial Network (GAN)-based framework designed to generate high-quality synthetic multi-omics profiles while preserving biological relationships. We evaluated Omics-GAN on three omics types (mRNA, miRNA, and DNA methylation) using the ROSMAP cohort for Alzheimer's disease (AD) and TCGA datasets for colon and liver cancer. A support vector machine (SVM) classifier with repeated 5-fold cross-validation demonstrated that synthetic datasets consistently improved prediction accuracy compared to original omics profiles. The AUC of SVM for mRNA improved from 0.72 to 0.74 in AD, and from 0.68 to 0.72 in liver cancer. Synthetic miRNA enhanced classification in colon cancer from 0.59 to 0.69, while synthetic methylation data improved performance in liver cancer from 0.64 to 0.71. Boxplot analyses confirmed that synthetic data preserved statistical distributions while reducing noise and outliers. Feature selection identified significant genes overlapping with original datasets and revealed additional candidates validated by GO and KEGG enrichment analyses. Finally, molecular docking highlighted potential drug repurposing candidates, including Nilotinib for AD, Atovaquone for liver cancer, and Tecovirimat for colon cancer. Omics-GAN enhances disease prediction, preserves biological fidelity, and accelerates biomarker and drug discovery, offering a scalable strategy for precision medicine applications.
Paper Structure (20 sections, 9 equations, 6 figures, 5 tables)

This paper contains 20 sections, 9 equations, 6 figures, 5 tables.

Figures (6)

  • Figure 1: Predicting the Alzheimer's Disease phenotype utilizing multi-omics datasets through Generative Adversarial Network techniques.
  • Figure 2: Prediction results of ROSMAP and TCGA (Liver and Colon cancer) datasets using validation samples. AUC of the prediction results using validation samples of synthetic Omics1, Omics2, and Omics3 with random network for $k = [1, 2, 3, 4, 5]$, where (A) illustrates the ROSMAP dataset with 245 and 351 subset AD patients, (B) illustrates liver cancer, and (C) illustrates colon cancer. Updated $k^*$ with the best validation AUC is selected as the final synthetic data for each omics profile.
  • Figure 3: Prediction results for AD and cancer patients using original, synthetic, and random networks for omics expression in ROSMAP and TCGA (Liver and Colon cancer) datasets. A indicates the prediction results for ROSMAP data. B indicates the prediction results for Liver cancer data, and C indicates the prediction results for Colon cancer data.
  • Figure 4: Significant features of the original and synthetic datasets are visualized using a Venn diagram, where (A) represents the ROSMAP dataset for AD patients, (B) represents the TCGA liver cancer dataset, and (C) represents the TCGA colon cancer dataset.
  • Figure 5: Image of binding affinities based on the top-ordered 30 drug agents against the ordered target proteins, where red colors indicated the strong binding affinities. (A) represents the binding score for AD patients, (B) represents the binding score for liver cancer, and (C) represents the binding score for colon cancer.
  • ...and 1 more figures