Table of Contents
Fetching ...

Accelerating Moment Tensor Potentials through Post-Training Pruning

Zijian Meng, Karim Zongo, Matthew Thoms, Ryan Eric Grant, Laurent Karim Béland

TL;DR

This work introduces a post-training, cost-aware pruning strategy that removes expensive basis functions with minimal loss of accuracy in Moment Tensor Potentials and yields models up to seven times faster than standard MTPs.

Abstract

Moment Tensor Potentials (MTPs) are machine-learning interatomic potentials whose basis functions are typically selected using a level-based scheme that is data-agnostic. We introduce a post-training, cost-aware pruning strategy that removes expensive basis functions with minimal loss of accuracy. Applied to nickel and silicon-oxygen systems, it yields models up to seven times faster than standard MTPs. The method requires no new data and remains fully compatible with current MTP implementations.

Accelerating Moment Tensor Potentials through Post-Training Pruning

TL;DR

This work introduces a post-training, cost-aware pruning strategy that removes expensive basis functions with minimal loss of accuracy in Moment Tensor Potentials and yields models up to seven times faster than standard MTPs.

Abstract

Moment Tensor Potentials (MTPs) are machine-learning interatomic potentials whose basis functions are typically selected using a level-based scheme that is data-agnostic. We introduce a post-training, cost-aware pruning strategy that removes expensive basis functions with minimal loss of accuracy. Applied to nickel and silicon-oxygen systems, it yields models up to seven times faster than standard MTPs. The method requires no new data and remains fully compatible with current MTP implementations.
Paper Structure (1 section, 5 figures)

This paper contains 1 section, 5 figures.

Figures (5)

  • Figure 1: Left: Nickel. Right: Silicon–oxygen. Cost–accuracy Pareto fronts obtained from the pruning process using NSGA-II and MOEA/D, each normalized to the base potential (MTP level 28). Twelve representative potentials, labeled A–L, were selected to span the Pareto fronts and illustrate the range of achievable cost–accuracy tradeoffs. NSGA-II consistently identified models with lower computational cost and higher accuracy compared with MOEA/D.
  • Figure 2: Top: Nickel. Bottom: Silicon–oxygen. Cost–accuracy comparison between the original level-based MTPs and the pruned MTPs. Panels show (left) training loss, (middle) per-atom energy (MAE), and (right) force MAE. “Inherited” and “Random” denote the two initialization methods for fitting the pruned potentials. Computational costs were measured on a single AMD EPYC 9654 core using systems of 2048 FCC Ni atoms and 1944 $\alpha$-quartz atoms. Representative examples with comparable or improved accuracy show speedups of 3.8$\times$ (Ni, model E) and 7.0$\times$ (Si–O, model F) relative to level-based MTPs.
  • Figure 3: Nickel MTP errors. Panels show, from top to bottom: lattice parameter; elastic constants $C_{11}$, $C_{12}$, and $C_{44}$ at 300 K; unstable stacking fault energy along $\langle 112 \rangle$ on the $\{111\}$ plane; $\Sigma3(110)$ tilt grain boundary formation energy; and formation and migration energies for monovacancies and $\langle 110 \rangle$ dumbbell self-interstitials. Pruned potentials exhibit similar error to the level-28 MTP, indicating that pruning primarily removes inefficient basis functions without greatly altering the fitted behavior. Color luminance indicates models within each group ordered by increasing cost.
  • Figure 4: Top to bottom: silicon lattice parameter, $\alpha$-quartz $a$, and $\alpha$-quartz $c$. Errors relative to DFT at 0 K are shown for the original level-based MTPs, the pruned models, and the reference level-26 potential from Zongo et al.zongo2024unified. The reference potential was trained with different hyperparameters and initialization settings from those used here. Pruned models reproduce these equilibrium properties with comparable accuracy while achieving substantial reductions in computational cost. Color luminance indicates models within each group ordered by increasing cost.
  • Figure 5: Top to bottom: liquid Si, amorphous Si, and liquid SiO$_2$. RDF (left) and ADF (right) are shown for the original level-based MTPs and the pruned models. DFT results and the reference level-26 MTP from Zongo et al.zongo2024unified are included for comparison. Experimental data is from Laaziri et al.laaziri1999high. Insets highlight regions of higher variability. Pruned models reproduce the main DFT and experimental features across all systems while substantially lowering computational cost. Color luminance indicates models within each group ordered by increasing cost; in silica, white denotes unstable potentials.