Table of Contents
Fetching ...

Boltzmann Graph Ensemble Embeddings for Aptamer Libraries

Starlika Bauskar, Jade Jiao, Narayanan Kannan, Alexander Kimm, Justin M. Baker, Matthew J. Tyler, Andrea L. Bertozzi, Anne M. Andrews

TL;DR

This work tackles SELEX-induced biases that obscure true aptamer binding by modeling aptamer secondary-structure ensembles as Boltzmann-weighted ERGMs. It defines two motifs, faces and rooted neighborhoods, and derives ensemble-expected fingerprints (ensemble fingerprints) that summarize motif occurrences under the Boltzmann distribution $p_{S,beta}(G) = \frac{e^{-beta E(G)}}{Z_S(beta)}$. By applying these chemistry-informed fingerprints to SELEX data, the authors demonstrate robust community detection and subgraph-level explainability, enabling identification of low-abundance candidates that warrant experimental follow-up. The approach offers a principled framework for anomaly detection and rational candidate prioritization in aptamer discovery, while acknowledging limitations such as exclusion of pseudoknots and fixed thermodynamic parameters; future directions include learning parameters, exploring multi-temperature ensembles, and experimental validation.

Abstract

Machine-learning methods in biochemistry commonly represent molecules as graphs of pairwise intermolecular interactions for property and structure predictions. Most methods operate on a single graph, typically the minimal free energy (MFE) structure, for low-energy ensembles (conformations) representative of structures at thermodynamic equilibrium. We introduce a thermodynamically parameterized exponential-family random graph (ERGM) embedding that models molecules as Boltzmann-weighted ensembles of interaction graphs. We evaluate this embedding on SELEX datasets, where experimental biases (e.g., PCR amplification or sequencing noise) can obscure true aptamer-ligand affinity, producing anomalous candidates whose observed abundance diverges from their actual binding strength. We show that the proposed embedding enables robust community detection and subgraph-level explanations for aptamer ligand affinity, even in the presence of biased observations. This approach may be used to identify low-abundance aptamer candidates for further experimental evaluation.

Boltzmann Graph Ensemble Embeddings for Aptamer Libraries

TL;DR

This work tackles SELEX-induced biases that obscure true aptamer binding by modeling aptamer secondary-structure ensembles as Boltzmann-weighted ERGMs. It defines two motifs, faces and rooted neighborhoods, and derives ensemble-expected fingerprints (ensemble fingerprints) that summarize motif occurrences under the Boltzmann distribution . By applying these chemistry-informed fingerprints to SELEX data, the authors demonstrate robust community detection and subgraph-level explainability, enabling identification of low-abundance candidates that warrant experimental follow-up. The approach offers a principled framework for anomaly detection and rational candidate prioritization in aptamer discovery, while acknowledging limitations such as exclusion of pseudoknots and fixed thermodynamic parameters; future directions include learning parameters, exploring multi-temperature ensembles, and experimental validation.

Abstract

Machine-learning methods in biochemistry commonly represent molecules as graphs of pairwise intermolecular interactions for property and structure predictions. Most methods operate on a single graph, typically the minimal free energy (MFE) structure, for low-energy ensembles (conformations) representative of structures at thermodynamic equilibrium. We introduce a thermodynamically parameterized exponential-family random graph (ERGM) embedding that models molecules as Boltzmann-weighted ensembles of interaction graphs. We evaluate this embedding on SELEX datasets, where experimental biases (e.g., PCR amplification or sequencing noise) can obscure true aptamer-ligand affinity, producing anomalous candidates whose observed abundance diverges from their actual binding strength. We show that the proposed embedding enables robust community detection and subgraph-level explanations for aptamer ligand affinity, even in the presence of biased observations. This approach may be used to identify low-abundance aptamer candidates for further experimental evaluation.
Paper Structure (16 sections, 10 equations, 6 figures, 1 table)

This paper contains 16 sections, 10 equations, 6 figures, 1 table.

Figures (6)

  • Figure 1: Chemically-informed aptamer embedding via secondary structure ensembles. From a sequence of nucleotides, we identify the distribution of secondary structure graphs. Each graph is mapped to a bag of face vector (defined below). Our final features are expected bag of faces vectors defined as sums weighted by the probability distribution.
  • Figure 2: Illustration of the five standard aptamer face types: hairpin, stack, internal loop, bulge and multibranch. Each face has a defining edge $(i,j)\in{\mathcal{E}}_\text{pair}$ shown by a black arc. The path between $(i,j)$ is represented by a dashed line and any nested regions bounded by an edge in ${\mathcal{E}}_\text{pair}$ are illustrated by blue arcs. In addition to type, each face has an associated energy depending on its nucleotide makeup. We consider (type, energy) pairs as subgraph motifs, as presented in gmfold2025.
  • Figure 3: Selective pressure versus count per million (center) with the histograms for normalized count (top) and selective pressure (right). Low-count high-pressure anomalies are marked in green and high-count low-pressure anomalies are marked in red.
  • Figure 4: Two-dimensional t-SNE embedding after applying NMF with 25 topics. Points are colored by cluster using spectral clustering to identify 35 clusters. Green X's are LC-HP aptamers, red X's are HC-LP aptamers, and two robust neighborhoods are circled in black.
  • Figure 5: Two-dimensional t-SNE embedding on negatively correlated features. Points are colored by cluster using spectral clustering to identify 25 clusters. Circled in black are identified anomalous clusters. Red X's are over-valued anomalies (HC-LP) and green X's are tested good binders. The structures of two captured over-valued anomalies are shown.
  • ...and 1 more figures