Table of Contents
Fetching ...

A Standardized Benchmark for Machine-Learned Molecular Dynamics using Weighted Ensemble Sampling

Alexander Aghili, Andy Bruce, Daniel Sabo, Sanya Murdeshwar, Kevin Bachelor, Ionut Mistreanu, Ashwin Lokapally, Razvan Marinescu

TL;DR

This work tackles the lack of standardized validation for MD and ML-driven dynamics by introducing a modular, WESTPA-based benchmarking framework that uses $TICA$-based progress coordinates to efficiently explore protein conformational space. It provides a ground-truth dataset of nine diverse proteins and a 19-metric evaluation suite, enabling reproducible, rigorous comparisons across classical and machine-learned MD methods. The authors demonstrate the framework by comparing an explicitly solvated ground truth with an implicit-solvent WE model and with CGSchNet-based ML dynamics, highlighting the ability to distinguish well-trained models from under-trained ones and to quantify both global structure and slow kinetics. By making the benchmark open-source and extensible, this work lays the groundwork for community-driven standardization in MD method validation with potential to accelerate development of reliable, scalable simulation tools in computational biophysics.

Abstract

The rapid evolution of molecular dynamics (MD) methods, including machine-learned dynamics, has outpaced the development of standardized tools for method validation. Objective comparison between simulation approaches is often hindered by inconsistent evaluation metrics, insufficient sampling of rare conformational states, and the absence of reproducible benchmarks. To address these challenges, we introduce a modular benchmarking framework that systematically evaluates protein MD methods using enhanced sampling analysis. Our approach uses weighted ensemble (WE) sampling via The Weighted Ensemble Simulation Toolkit with Parallelization and Analysis (WESTPA), based on progress coordinates derived from Time-lagged Independent Component Analysis (TICA), enabling fast and efficient exploration of protein conformational space. The framework includes a flexible, lightweight propagator interface that supports arbitrary simulation engines, allowing both classical force fields and machine learning-based models. Additionally, the framework offers a comprehensive evaluation suite capable of computing more than 19 different metrics and visualizations across a variety of domains. We further contribute a dataset of nine diverse proteins, ranging from 10 to 224 residues, that span a variety of folding complexities and topologies. Each protein has been extensively simulated at 300K for one million MD steps per starting point (4 ns). To demonstrate the utility of our framework, we perform validation tests using classic MD simulations with implicit solvent and compare protein conformational sampling using a fully trained versus under-trained CGSchNet model. By standardizing evaluation protocols and enabling direct, reproducible comparisons across MD approaches, our open-source platform lays the groundwork for consistent, rigorous benchmarking across the molecular simulation community.

A Standardized Benchmark for Machine-Learned Molecular Dynamics using Weighted Ensemble Sampling

TL;DR

This work tackles the lack of standardized validation for MD and ML-driven dynamics by introducing a modular, WESTPA-based benchmarking framework that uses -based progress coordinates to efficiently explore protein conformational space. It provides a ground-truth dataset of nine diverse proteins and a 19-metric evaluation suite, enabling reproducible, rigorous comparisons across classical and machine-learned MD methods. The authors demonstrate the framework by comparing an explicitly solvated ground truth with an implicit-solvent WE model and with CGSchNet-based ML dynamics, highlighting the ability to distinguish well-trained models from under-trained ones and to quantify both global structure and slow kinetics. By making the benchmark open-source and extensible, this work lays the groundwork for community-driven standardization in MD method validation with potential to accelerate development of reliable, scalable simulation tools in computational biophysics.

Abstract

The rapid evolution of molecular dynamics (MD) methods, including machine-learned dynamics, has outpaced the development of standardized tools for method validation. Objective comparison between simulation approaches is often hindered by inconsistent evaluation metrics, insufficient sampling of rare conformational states, and the absence of reproducible benchmarks. To address these challenges, we introduce a modular benchmarking framework that systematically evaluates protein MD methods using enhanced sampling analysis. Our approach uses weighted ensemble (WE) sampling via The Weighted Ensemble Simulation Toolkit with Parallelization and Analysis (WESTPA), based on progress coordinates derived from Time-lagged Independent Component Analysis (TICA), enabling fast and efficient exploration of protein conformational space. The framework includes a flexible, lightweight propagator interface that supports arbitrary simulation engines, allowing both classical force fields and machine learning-based models. Additionally, the framework offers a comprehensive evaluation suite capable of computing more than 19 different metrics and visualizations across a variety of domains. We further contribute a dataset of nine diverse proteins, ranging from 10 to 224 residues, that span a variety of folding complexities and topologies. Each protein has been extensively simulated at 300K for one million MD steps per starting point (4 ns). To demonstrate the utility of our framework, we perform validation tests using classic MD simulations with implicit solvent and compare protein conformational sampling using a fully trained versus under-trained CGSchNet model. By standardizing evaluation protocols and enabling direct, reproducible comparisons across MD approaches, our open-source platform lays the groundwork for consistent, rigorous benchmarking across the molecular simulation community.
Paper Structure (29 sections, 7 equations, 10 figures, 9 tables)

This paper contains 29 sections, 7 equations, 10 figures, 9 tables.

Figures (10)

  • Figure 1: Architecture of our weighted ensemble (WE) benchmarking framework.
  • Figure 2: Ground truth trajectories of Protein G in TICA 0/1 space with a 10% sampling of starting points, colored by starting point number. The legend provides the starting point number using R# where # is the starting point associated with the color (divided by 10). This is not representative of the explored space, but shows how little the dynamics will deviate from the original point.
  • Figure 3: Complete set of ground truth points for Protein G in TICA 0/1 space with a stride of 100. The color of each point is the value of the 3rd TICA component.
  • Figure 4: Conformations randomly chosen based on TICA proximity, visualized along side their location in TICA space. A: Many overlayed visualizations of close conformations with close TICA proximity. B: Visualization of randomly chosen conformation located far away in TICA space. C: Projection of all conformations visualized into TIC 0, and TIC 1 space.
  • Figure 5: Construction of the MSM starts with (left) the Log probability of transition matrix between the discrete bins, then (middle) the equilibrium distribution using the MSM vs the raw quantities from the WESTPA run, and finally (right) constructing the two KDEs (Green is the SSMSM, Blue is the Raw Data).
  • ...and 5 more figures