Table of Contents
Fetching ...

Functional and parametric identifiability for universal differential equations applied to chemical reaction networks

Torkel E Loman, Ruth E Baker

TL;DR

This work shows that universal differential equations (UDEs)—neural networks embedded in mechanistic CRN-driven ODEs—can maintain parametric identifiability while enabling learning of unknown dynamics. Identifiability is split into parametric (likelihood-based) and functional (ensemble-based) components, with parametric identifiability assessed via profile likelihood and functional identifiability via ensembles of fitted functions. Across diverse CRN scenarios, replacing known functional forms with small neural networks yields little loss in parametric identifiability, and identifiability can be further improved by enforcing neural network constraints such as monotonicity and bounds. The results highlight the robustness of CRN-UDEs to misspecification and their potential for interpretable, data-driven discovery in biology, chemistry, pharmacology, and epidemiology.

Abstract

Mathematical modelling has traditionally relied on detailed system knowledge to construct mechanistic models. However, the advent of large-scale data collection and advances in machine learning have led to an increasing use of data-driven approaches. Recently, hybrid models have emerged that combine both paradigms: well-understood system components are modelled mechanistically, while unknown parts are inferred from data. Here, we focus on one such class: universal differential equations (UDEs), where neural networks are embedded within differential equations to approximate unknown dynamics. When fitted to data, these networks act as universal function approximators, learning missing functional components. In this work, we note that UDE identifiability, i.e. our ability to identify true system properties, can be split into parametric and functional identifiability (assessing identifiability for the mechanistic and data-driven model parts, respectively). Next, we investigate how UDE properties, such as neural network numbers and constraints, affect parametric and functional identifiability. Notably, we show that across a wide range of models, the generalisation of a fully mechanistic model to a UDE has little impact on the mechanistic components' parametric identifiability. Finally, we note that hybrid modelling through the fitting of unknown functions (as achieved by UDEs) is particularly well-suited to chemical reaction network (CRN) modelling. Here, CRNs are used in fields ranging from systems biology, chemistry, and pharmacology to epidemiology and population dynamics, making them highly relevant for study. By showcasing how CRN-based UDE models can be highly interpretable, we demonstrate that this hybrid approach is a promising avenue for future applications.

Functional and parametric identifiability for universal differential equations applied to chemical reaction networks

TL;DR

This work shows that universal differential equations (UDEs)—neural networks embedded in mechanistic CRN-driven ODEs—can maintain parametric identifiability while enabling learning of unknown dynamics. Identifiability is split into parametric (likelihood-based) and functional (ensemble-based) components, with parametric identifiability assessed via profile likelihood and functional identifiability via ensembles of fitted functions. Across diverse CRN scenarios, replacing known functional forms with small neural networks yields little loss in parametric identifiability, and identifiability can be further improved by enforcing neural network constraints such as monotonicity and bounds. The results highlight the robustness of CRN-UDEs to misspecification and their potential for interpretable, data-driven discovery in biology, chemistry, pharmacology, and epidemiology.

Abstract

Mathematical modelling has traditionally relied on detailed system knowledge to construct mechanistic models. However, the advent of large-scale data collection and advances in machine learning have led to an increasing use of data-driven approaches. Recently, hybrid models have emerged that combine both paradigms: well-understood system components are modelled mechanistically, while unknown parts are inferred from data. Here, we focus on one such class: universal differential equations (UDEs), where neural networks are embedded within differential equations to approximate unknown dynamics. When fitted to data, these networks act as universal function approximators, learning missing functional components. In this work, we note that UDE identifiability, i.e. our ability to identify true system properties, can be split into parametric and functional identifiability (assessing identifiability for the mechanistic and data-driven model parts, respectively). Next, we investigate how UDE properties, such as neural network numbers and constraints, affect parametric and functional identifiability. Notably, we show that across a wide range of models, the generalisation of a fully mechanistic model to a UDE has little impact on the mechanistic components' parametric identifiability. Finally, we note that hybrid modelling through the fitting of unknown functions (as achieved by UDEs) is particularly well-suited to chemical reaction network (CRN) modelling. Here, CRNs are used in fields ranging from systems biology, chemistry, and pharmacology to epidemiology and population dynamics, making them highly relevant for study. By showcasing how CRN-based UDE models can be highly interpretable, we demonstrate that this hybrid approach is a promising avenue for future applications.
Paper Structure (28 sections, 29 equations, 19 figures)

This paper contains 28 sections, 29 equations, 19 figures.

Figures (19)

  • Figure 1: Profile likelihood analysis can be used to assess parameter identifiability. For a parameter fitting optimisation problem, a profile likelihood diagram can be computed for each parameter kreutz_profile_2013. The x-axis denotes the parameter's range of potential values. For each parameter value, the y-axis denotes the best loss value that can be achieved while keeping the parameter fixed at that value (and all remaining parameters are free). Naively, the likelihood profile can be computed by solving the parameter fitting optimisation problem along a grid of parameter values, however, more efficient approaches exist noauthor_confidence_nodate. By considering its likelihood profile, a parameter's identifiability can be assessed. (A) A profile exhibiting a distinct peak around a single value suggests identifiability. Here, only parameter values around the peak generate a good model-to-data fit, indicating that the peak corresponds to the parameter's true value. (B) A flat profile indicates that the parameter can be varied arbitrarily without affecting the fit. This corresponds to structural non-identifiability. (C) A likelihood profile with a plateau suggests that the parameter can attain a wide range of values. This corresponds to practical non-identifiability. (A-C) For maximum likelihood optimisation, if the profile is shifted in the y-direction so that the maximum likelihood estimate occurs at $0$, a line can be drawn at $y \approx 1.92$. The intersection of this line with the profile corresponds to the parameter's $95\%$ confidence interval. A narrow confidence interval corresponds to high identifiability.
  • Figure 1: Practical identifiability analysis of the extended self-activation loop: Repeat 1. This figure is an exact repeat of Figure \ref{['figure:extended_sal_pract_ident']}. The only difference is that the data have been resampled (but using the same distribution). The repeat in this figure replicates results in the main figure.
  • Figure 2: The simple self-activation loop UDE exhibits non-identifiability. (A) A ground truth simulation (solid line), from which we have generated synthetic data (dots), to which we fit three models; $A(X) = 0.6\frac{X^3}{X^3 + 0.3^3}$ (dashed line), $A(X) = v\frac{X^n}{X^n + K^n}$ (dotted line), and $A(X) = U(X;\mathbf{\theta})$ (dash-dotted line). All models fit the data well (here, all sets of lines overlap closely). (B) Likelihood profiles for the parameter $d$ for each model. The 95% confidence interval is marked with a horizontal dashed line and $d$'s true value with a vertical dotted line. The UDE yields a flat profile, suggesting non-identifiability. (C) The ensemble of functions $A$ that was successfully fitted for the parameterised Hill function model (purple lines) and UDE (green line), as well as the ground truth function (dashed yellow line). Here, the UDE yields a wide range of functions $A$ that fit the data, suggesting that $A$ is non-identifiable.
  • Figure 2: Practical identifiability analysis of the extended self-activation loop: Repeat 2. This figure is an exact repeat of Figure \ref{['figure:extended_sal_pract_ident']}. The only difference is that the data have been resampled (but using the same distribution). The repeat in this figure replicates results in the main figure.
  • Figure 3: The extended self-activation loop may exhibit practical non-identifiability. We fit three models ($(A(Y) = 1.1\frac{Y^3}{Y^3 + 2^3}$, $v\frac{Y^n}{Y^n + K^n}$, and $U(Y,\mathbf{\theta})$) to three different data sets: $X$ measured at low noise (A,D,G), $X$ and $Y$ measured at low noise (B,E,H), and $X$ and $Y$ measured at high noise (C,F,I). (A-C) Model ground truth simulations (solid lines), measured data (dots), and simulations of the fitted models (dashed and/or dotted lines). $X$ is displayed in shades of blue and $Y$ in shades of red. In all cases, the model replicates the measured data. However, when $X$ only is measured, the parameterised Hill and UDE models yield erroneous predictions for $Y$. (D,E,F) Likelihood profiles for $d$. $d$ is highly identifiable, except for the two cases which yielded erroneous predictions in A. In the remaining cases, the parameterised Hill and UDE models yields similar levels of identifiability, and both pinpoint $d$ well even in the high-noise case. (G,H,I) For each dataset, we plot the forms of fitted $A$ functions for the parameterised Hill (purple lines) and UDE (green lines) as compared to the ground-truth (dashed yellow lines). The two models exhibit functional non-identifiability in the same cases where they exhibit parametric non-identifiability and yield incorrect predictions of $Y$. In H and I, however, both models pinpoint ground-truth $A$ well.
  • ...and 14 more figures