Transforming Multi-Omics Integration with GANs: Applications in Alzheimer's and Cancer
Md Selim Reza, Sabrin Afroz, Mostafizer Rahman, Md Ashad Alam
TL;DR
This work tackles the challenges of integrating heterogeneous multi-omics data under limited samples and high noise by introducing Omics-GAN, a three-WGAN framework that generates high-quality synthetic profiles for mRNA, DNA methylation, and miRNA while preserving inter-omics relationships. Applied to ROSMAP (Alzheimer’s) and TCGA liver/colon cancer, Omics-GAN consistently improved disease outcome predictions (e.g., higher AUC in SVM classifiers) and revealed biologically meaningful features that persisted across original and synthetic data, validated by GO and KEGG enrichment analyses. The study further demonstrates the value of incorporating interaction networks, showing that true biological networks enhance predictive gains over random networks, and leverages molecular docking to propose drug repurposing candidates (Nilotinib for AD, Atovaquone for liver cancer, Tecovirimat for colon cancer). Collectively, Omics-GAN advances accurate biomarker discovery and therapeutic prioritization by generating faithful synthetic multi-omics data, offering a scalable path for precision medicine.
Abstract
Multi-omics data integration is crucial for understanding complex diseases, yet limited sample sizes, noise, and heterogeneity often reduce predictive power. To address these challenges, we introduce Omics-GAN, a Generative Adversarial Network (GAN)-based framework designed to generate high-quality synthetic multi-omics profiles while preserving biological relationships. We evaluated Omics-GAN on three omics types (mRNA, miRNA, and DNA methylation) using the ROSMAP cohort for Alzheimer's disease (AD) and TCGA datasets for colon and liver cancer. A support vector machine (SVM) classifier with repeated 5-fold cross-validation demonstrated that synthetic datasets consistently improved prediction accuracy compared to original omics profiles. The AUC of SVM for mRNA improved from 0.72 to 0.74 in AD, and from 0.68 to 0.72 in liver cancer. Synthetic miRNA enhanced classification in colon cancer from 0.59 to 0.69, while synthetic methylation data improved performance in liver cancer from 0.64 to 0.71. Boxplot analyses confirmed that synthetic data preserved statistical distributions while reducing noise and outliers. Feature selection identified significant genes overlapping with original datasets and revealed additional candidates validated by GO and KEGG enrichment analyses. Finally, molecular docking highlighted potential drug repurposing candidates, including Nilotinib for AD, Atovaquone for liver cancer, and Tecovirimat for colon cancer. Omics-GAN enhances disease prediction, preserves biological fidelity, and accelerates biomarker and drug discovery, offering a scalable strategy for precision medicine applications.
