Table of Contents
Fetching ...

Kernel-Based Evaluation of Conditional Biological Sequence Models

Pierre Glaser, Steffanie Paul, Alissa M. Hummer, Charlotte M. Deane, Debora S. Marks, Alan N. Amin

TL;DR

This work develops kernel-based tools for evaluating conditional sequence models, introducing the Augmented Conditional MMD ($\operatorname{ACMMD}$) as an absolute measure of how well a conditional sequence model $Q_{|}$ matches the true conditional $\mathbb{P}_{|}$. It also defines $\operatorname{ACMMD--Rel}$ to assess model reliability and provides unbiased estimators and hypothesis tests (via wild bootstrap) that control Type I error in finite samples. The authors prove that ACMMD discriminates any mismatch under universal kernels and demonstrate its utility through synthetic examples and a detailed case study of ProteinMPNN on the CATH dataset, including temperature-tuning insights and cross-family variability. Overall, the framework offers a principled, statistically sound approach to quantify both accuracy and reliability of conditional sequence models with potential impact on protein design and beyond.

Abstract

We propose a set of kernel-based tools to evaluate the designs and tune the hyperparameters of conditional sequence models, with a focus on problems in computational biology. The backbone of our tools is a new measure of discrepancy between the true conditional distribution and the model's estimate, called the Augmented Conditional Maximum Mean Discrepancy (ACMMD). Provided that the model can be sampled from, the ACMMD can be estimated unbiasedly from data to quantify absolute model fit, integrated within hypothesis tests, and used to evaluate model reliability. We demonstrate the utility of our approach by analyzing a popular protein design model, ProteinMPNN. We are able to reject the hypothesis that ProteinMPNN fits its data for various protein families, and tune the model's temperature hyperparameter to achieve a better fit.

Kernel-Based Evaluation of Conditional Biological Sequence Models

TL;DR

This work develops kernel-based tools for evaluating conditional sequence models, introducing the Augmented Conditional MMD () as an absolute measure of how well a conditional sequence model matches the true conditional . It also defines to assess model reliability and provides unbiased estimators and hypothesis tests (via wild bootstrap) that control Type I error in finite samples. The authors prove that ACMMD discriminates any mismatch under universal kernels and demonstrate its utility through synthetic examples and a detailed case study of ProteinMPNN on the CATH dataset, including temperature-tuning insights and cross-family variability. Overall, the framework offers a principled, statistically sound approach to quantify both accuracy and reliability of conditional sequence models with potential impact on protein design and beyond.

Abstract

We propose a set of kernel-based tools to evaluate the designs and tune the hyperparameters of conditional sequence models, with a focus on problems in computational biology. The backbone of our tools is a new measure of discrepancy between the true conditional distribution and the model's estimate, called the Augmented Conditional Maximum Mean Discrepancy (ACMMD). Provided that the model can be sampled from, the ACMMD can be estimated unbiasedly from data to quantify absolute model fit, integrated within hypothesis tests, and used to evaluate model reliability. We demonstrate the utility of our approach by analyzing a popular protein design model, ProteinMPNN. We are able to reject the hypothesis that ProteinMPNN fits its data for various protein families, and tune the model's temperature hyperparameter to achieve a better fit.
Paper Structure (55 sections, 14 theorems, 100 equations, 8 figures, 3 algorithms)

This paper contains 55 sections, 14 theorems, 100 equations, 8 figures, 3 algorithms.

Key Result

Lemma 3.2

Under mild integrability conditions, we have: Where $\mu_{\mathbb{ P}_{|}}$ and $\mu_{Q_{|}}$ are the conditional mean embeddings park2020measure of $\mathbb P_{|}$ and $Q_{|}$, $K_{\mathcal{X}}(x, x') \coloneqq k_{\mathcal{X}}(x, x') I_{\mathcal{H}_{\mathcal{Y}}}$ (here, $I_{\mathcal{H}_{\mathcal{Y}}}$ the identity operator) is an operator-va

Figures (8)

  • Figure 1: Left panel: ACMMD estimates for a fixed shift value $\Delta p = 0.25$ and various number of samples in the synthetic example of \ref{['sec:toy-synthetic']}. The analytic ACMMD value is given by the horizontal line. Right panel: ACMMD test average rejection rate for various number of samples and shifts in the same setting.
  • Figure 2: Values of $\widehat{ \operatorname{ACMMD} }\newline^2$, (left) and of the average rejection rate of the ACMMD test (right) in the setting described in \ref{['sec:discriminative']}. Each line corresponds to a different value for $\delta T$.
  • Figure 3: Values of $\widehat{ \operatorname{ACMMD} }\operatorname{--Rel}\newline^2$, (left) and of the average rejection rate of the ACMMD--Rel test (right) in the setting described in \ref{['sec:discriminative']}. Each line corresponds to a different value for $\delta T$.
  • Figure 4: Evolution of $\widehat{ \operatorname{ACMMD} }\newline^2$ (left) and $\widehat{ \operatorname{ACMMD}}\operatorname{--Rel}\newline^2$ (right) between a pre-trained ProteinMPNN model and the CATH S60 reference dataset, for varying temperature values
  • Figure 5: Estimated value $\widehat{ \operatorname{ACMMD} }\newline^2$ between ProteinMPNN and the CATH S60 reference dataset on a subset of 10 superfamilies for two different temperatures $T=1.0$ and $T=0.1$.
  • ...and 3 more figures

Theorems & Definitions (23)

  • Definition 3.1: Augmented Conditional MMD
  • Lemma 3.2
  • Lemma 3.3
  • Lemma 3.4
  • Definition 4.1: ACMMD for Reliability
  • Proposition 4.2
  • Proposition 4.3
  • Proposition 4.4
  • Lemma : Complete form of Lemma \ref{['lemma:cgof_mmd']}
  • proof
  • ...and 13 more