Table of Contents
Fetching ...

Stoichiometrically-informed symbolic regression for extracting chemical reaction mechanisms from data

Manuel Palma Banos, Joel D. Kress, Rigoberto Hernandez, Galen T. Craven

TL;DR

This work tackles automatic extraction of chemical reaction mechanisms from time-series concentration data by introducing Stoichiometrically-Informed Symbolic Regression (SISR). SISR searches a space of stoichiometry-constrained reactions with a genetic algorithm and fits rate constants in derivative space, balancing data fit with model simplicity via a multiobjective Pareto approach that uses $ ext{L}_c$ and a complexity metric. Across paradigmatic systems including Sequential Linear, Lotka–Volterra, nonlinear fast/slow dynamics, Michaelis–Menten, and glucose oxidation, SISR exactly recovers the true mechanisms and rate constants, even in noise and when intermediates are hidden, and it outperforms SINDy-like methods that lack stoichiometric constraints. The approach yields interpretable, sparse kinetic equations and can forecast unseen data regions, highlighting its potential for data-driven chemical mechanism discovery. Limitations include purely numerical rate constants under fixed conditions and no current treatment of nonequilibrium or parameter-varying environments; future work aims to extend to nonequilibrium regimes and improve computational efficiency and experimental validation.

Abstract

A data-driven computational method is introduced to extract chemical reaction mechanisms from time series chemical concentration data. It is realized through the use of dynamic symbolic regression in which a sparse analytical form for a dynamical system is discoverable from the underlying data. We specifically develop the stoichiometrically-informed symbolic regression (SISR) method to address a standing challenge in complex chemical reaction networks: Given a time-series dataset of concentrations of several components, what is the mechanism and the associated rate constants? SISR finds the optimal mechanism, kinetic equations and rate constants by combining differential optimization with a genetic optimization approach that searches a symbolic space of possible reaction mechanisms. Use of SISR in several paradigmatic examples spanning linear and nonlinear reaction schemes results in excellent agreement between true and predicted mechanisms, including when the method is applied to noisy data. The advantages of a stoichiometrically-informed approach such as SISR to address reaction discovery is illustrated through comparison with the use of generic state-of-the-art data-driven approaches.

Stoichiometrically-informed symbolic regression for extracting chemical reaction mechanisms from data

TL;DR

This work tackles automatic extraction of chemical reaction mechanisms from time-series concentration data by introducing Stoichiometrically-Informed Symbolic Regression (SISR). SISR searches a space of stoichiometry-constrained reactions with a genetic algorithm and fits rate constants in derivative space, balancing data fit with model simplicity via a multiobjective Pareto approach that uses and a complexity metric. Across paradigmatic systems including Sequential Linear, Lotka–Volterra, nonlinear fast/slow dynamics, Michaelis–Menten, and glucose oxidation, SISR exactly recovers the true mechanisms and rate constants, even in noise and when intermediates are hidden, and it outperforms SINDy-like methods that lack stoichiometric constraints. The approach yields interpretable, sparse kinetic equations and can forecast unseen data regions, highlighting its potential for data-driven chemical mechanism discovery. Limitations include purely numerical rate constants under fixed conditions and no current treatment of nonequilibrium or parameter-varying environments; future work aims to extend to nonequilibrium regimes and improve computational efficiency and experimental validation.

Abstract

A data-driven computational method is introduced to extract chemical reaction mechanisms from time series chemical concentration data. It is realized through the use of dynamic symbolic regression in which a sparse analytical form for a dynamical system is discoverable from the underlying data. We specifically develop the stoichiometrically-informed symbolic regression (SISR) method to address a standing challenge in complex chemical reaction networks: Given a time-series dataset of concentrations of several components, what is the mechanism and the associated rate constants? SISR finds the optimal mechanism, kinetic equations and rate constants by combining differential optimization with a genetic optimization approach that searches a symbolic space of possible reaction mechanisms. Use of SISR in several paradigmatic examples spanning linear and nonlinear reaction schemes results in excellent agreement between true and predicted mechanisms, including when the method is applied to noisy data. The advantages of a stoichiometrically-informed approach such as SISR to address reaction discovery is illustrated through comparison with the use of generic state-of-the-art data-driven approaches.
Paper Structure (20 sections, 51 equations, 15 figures, 2 tables)

This paper contains 20 sections, 51 equations, 15 figures, 2 tables.

Figures (15)

  • Figure 1: Schematic diagram showing the workflow for the developed SISR method.
  • Figure 2: Expression tree complexity analysis for the example mechanism shown in Eq. (\ref{['eq:example2']})
  • Figure 3: Time evolution of the (a) concentrations and (b) scaled derivatives---Eq. (\ref{['eq:ScaledDerivative']})--- of each species in the sequential linear mechanism given in Eq. (\ref{['eq:seqmech']}). The solid lines are the results of the SISR method and the corresponding black markers are a subset of the data used by SISR to extract the reaction mechanism and fit the rate constants. The concentrations are shown in units of millimolar and time is shown in units of seconds.
  • Figure 4: Minimum derivative error, min($\mathcal{L}_i$), as a function of generation using SISR to extract mechanisms from the ground truth data generated using the mechanism given in Eq. (\ref{['eq:seqmech']}). The $y$-axis is shown on a log scale. Each curve is the result of a different island in the SISR method, where each island contains mechanisms with the number of reactions $|\bold{M}|$ shown in the legend.
  • Figure 5: Complexity vs. concentration error $\mathcal{L}_\text{c}$ for the sequential linear mechanism. Each marker corresponds to the best mechanism as calculated using the derivative error $\mathcal{L}_\text{der}$ from the labeled island. The $y$-axis is shown on a log scale.
  • ...and 10 more figures