Table of Contents
Fetching ...

SimPoly: Simulation of Polymers with Machine Learning Force Fields Derived from First Principles

Gregor N. C. Simm, Jean Hélie, Hannes Schulz, Yicheng Chen, Guillem Simeon, Anna Kuzina, Ernesto Martinez-Baez, Piero Gasparotto, Gabriele Tocci, Chi Chen, Yatao Li, Lixue Cheng, Zun Wang, Bichlien H. Nguyen, Jake A. Smith, Lixin Sun

TL;DR

This work demonstrates that a first-principles-trained ML force field for polymers (Vivace) can accurately predict experimental bulk properties, including densities and glass transition temperatures, across a broad polymer spectrum. It introduces PolyArena as an experimental benchmark and PolyData as a comprehensive training corpus, enabling ab initio learning that generalizes to unseen polymers. The key innovations—a localized SE(3)-equivariant GNN with a lightweight three-body tensor product and a multi-cutoff strategy—deliver fast, scalable, and accurate simulations that outperform classical force fields and rival other MLFFs. The study highlights both the practical impact of MLFFs for polymer design and the remaining challenges, such as long-range electrostatics, while laying groundwork for an in silico design pipeline for next-generation polymers.

Abstract

Polymers are a versatile class of materials with widespread industrial applications. Advanced computational tools could revolutionize their design, but their complex, multi-scale nature poses significant modeling challenges. Conventional force fields often lack the accuracy and transferability required to capture the intricate interactions governing polymer behavior. Conversely, quantum-chemical methods are computationally prohibitive for the large systems and long timescales required to simulate relevant polymer phenomena. Here, we overcome these limitations with a machine learning force field (MLFF) approach. We demonstrate that macroscopic properties for a broad range of polymers can be predicted ab initio, without fitting to experimental data. Specifically, we develop a fast and scalable MLFF to accurately predict polymer densities, outperforming established classical force fields. Our MLFF also captures second-order phase transitions, enabling the prediction of glass transition temperatures. To accelerate progress in this domain, we introduce a benchmark of experimental bulk properties for 130 polymers and an accompanying quantum-chemical dataset. This work lays the foundation for a fully in silico design pipeline for next-generation polymeric materials.

SimPoly: Simulation of Polymers with Machine Learning Force Fields Derived from First Principles

TL;DR

This work demonstrates that a first-principles-trained ML force field for polymers (Vivace) can accurately predict experimental bulk properties, including densities and glass transition temperatures, across a broad polymer spectrum. It introduces PolyArena as an experimental benchmark and PolyData as a comprehensive training corpus, enabling ab initio learning that generalizes to unseen polymers. The key innovations—a localized SE(3)-equivariant GNN with a lightweight three-body tensor product and a multi-cutoff strategy—deliver fast, scalable, and accurate simulations that outperform classical force fields and rival other MLFFs. The study highlights both the practical impact of MLFFs for polymer design and the remaining challenges, such as long-range electrostatics, while laying groundwork for an in silico design pipeline for next-generation polymers.

Abstract

Polymers are a versatile class of materials with widespread industrial applications. Advanced computational tools could revolutionize their design, but their complex, multi-scale nature poses significant modeling challenges. Conventional force fields often lack the accuracy and transferability required to capture the intricate interactions governing polymer behavior. Conversely, quantum-chemical methods are computationally prohibitive for the large systems and long timescales required to simulate relevant polymer phenomena. Here, we overcome these limitations with a machine learning force field (MLFF) approach. We demonstrate that macroscopic properties for a broad range of polymers can be predicted ab initio, without fitting to experimental data. Specifically, we develop a fast and scalable MLFF to accurately predict polymer densities, outperforming established classical force fields. Our MLFF also captures second-order phase transitions, enabling the prediction of glass transition temperatures. To accelerate progress in this domain, we introduce a benchmark of experimental bulk properties for 130 polymers and an accompanying quantum-chemical dataset. This work lays the foundation for a fully in silico design pipeline for next-generation polymeric materials.
Paper Structure (33 sections, 33 equations, 16 figures, 7 tables, 2 algorithms)

This paper contains 33 sections, 33 equations, 16 figures, 7 tables, 2 algorithms.

Figures (16)

  • Figure 1: Prediction of experimental bulk properties for a wide range of polymers through MD simulations with a MLFF trained solely on ab initio data. a, We generate a targeted quantum-chemical dataset containing small, atomistic polymeric model systems covering the complex intra- and inter-molecular interactions characteristic of polymers called PolyData. b, We introduce Vivace, a fast and scalable, SE(3)-equivariant MLFF architecture optimized for large-scale MD simulations, and train it on PolyData, together with a collection of other public data sets. c, We run MD simulations driven by Vivace using large model systems of polymers to predict experimentally measured (volumetric mass) densities. We also observe second-order phase transitions, enabling the determination of glass transition temperatures, T$_\mathrm{g}$ s. To systematically evaluate the performance of MLFFs on this challenging task, we introduce PolyArena, a benchmark of experimental bulk properties for 130.0 polymers.
  • Figure 2: Overview of the experimental and computational data presented in this study. PolyArena (a--d) provides a benchmark for developing and systematically evaluating MLFFs on soft matter systems containing experimentally measured densities and glass transition temperatures. PolyData (e--j) is a collection of training datasets for polymeric systems, consisting of three subsets, PolyPack, PolyDiss, and PolyCrop. a, Number of polymers containing at least one atom of the respective element. b, Distribution of molecular weights of repeating units. c, Distribution of experimental (volumetric mass) densities under standard conditions. d, Distribution of experimental glass transition temperatures T$_\mathrm{g}$. e, Example structure from PolyPack, which contains structures with multiple polymer chains packed in periodic boxes at different densities. f, Distribution of the norm of the nuclear forces in PolyPack. g, Example structure from PolyDiss shown with one periodic image. The "offset" is the minimal interatomic distance across the periodic boundary. PolyDiss structures contain a single polymer chain in a periodic box of increasing size, highlighting the inter-chain interactions of polymers. h, Dissociation curve of a single polymer chain in a periodic box. i, Example structure from PolyCrop, which contains fragments of polymer chains in vacuum. j, Distribution of the longest dimension of the PolyCrop clusters. k, l, Repeating units of a selection of polymers from the PolyArena benchmark that are studied in more detail. Polymers that were part of the training data are shown in k, while those that were not are shown in l. End groups are chosen for structural stability, not synthetic feasibility.
  • Figure 3: With Vivace, trained solely on ab initio data, one can accurately predict experimental densities for a wide range of polymers. In each panel, we plot calculated vs experimental densities of polymers from the PolyArena benchmark under standard conditions for a different FF. Each marker corresponds to an MD simulation from which the density was calculated. A selected list of 10 polymers are indicated by green markers. Due to its slow simulation speed, UMA was only applied to these polymers. Not all polymers were simulated with the remaining models due to simulations failing (see main text for details). In a, circles ($\hbox{$\bigcirc$}$) and triangles ($\triangle$) indicate polymers seen and unseen during Vivace's training, respectively. The polymers PS (seen) and PCTFE (unseen) are annotated. In panels b, c, d and e we do not make this distinction and indicate all polymers with the same square marker ($\square$). MAEs (in ) are reported in each panel. The dashed lines indicate perfect agreement, and the shaded areas show a deviation of ± 0.05.
  • Figure 4: Calculations using different FFs of the glass transition temperatures of 10 polymers selected from the PolyArena benchmark. a, b, Density vs temperature curves obtained for two selected polymers using Vivace. Each point represents the average of three density simulations, with weights inversely proportional to their uncertainty. Error bars indicate the combined uncertainty. Examples of fits used to derive T$_\mathrm{g}^\mathrm{fit}$ are shown in blue. T$_\mathrm{g}^\mathrm{fit}$ and the experimental T$_\mathrm{g}$ are indicated by dashed and dotted vertical lines, respectively. The shaded area indicates the uncertainty of the fit, derived by bootstrapping. More density vs temperature curves are shown in Figs. \ref{['fig:tg_mlff']}, \ref{['fig:tg_mace']}, and \ref{['fig:tg_pcff']}. c, Absolute errors (and corresponding uncertainty) of calculated vs experimental T$_\mathrm{g}$'s for different FFs. For Vivace we indicate polymers that were not seen during training by a hatched pattern. In the case of PAN, the large uncertainty range of 211 is not plotted ($\dagger$). Asterisks ($\ast$) indicate failed MD simulations so that no T$_\mathrm{g}$ could be determined. MAEs (in ) are reported across all polymers for which T$_\mathrm{g}$ could be determined.
  • Figure 5: Fine-tuning Vivace improves its ability to describe inter-chain interactions. The figure shows selected dissociation curves from the PolyDiss test set, computed with four MLFFs and the density functional theory (DFT) reference method (black). Each panel plots the relative energy (in ) as a function of the distance (offset) between polymer chains (in ) (see Section \ref{['sec:poly_data']} for details). To account for differences in absolute energies between models, the curves are shifted vertically so that the energy at the largest separation is zero.
  • ...and 11 more figures