Boltzmann Graph Ensemble Embeddings for Aptamer Libraries
Starlika Bauskar, Jade Jiao, Narayanan Kannan, Alexander Kimm, Justin M. Baker, Matthew J. Tyler, Andrea L. Bertozzi, Anne M. Andrews
TL;DR
This work tackles SELEX-induced biases that obscure true aptamer binding by modeling aptamer secondary-structure ensembles as Boltzmann-weighted ERGMs. It defines two motifs, faces and rooted neighborhoods, and derives ensemble-expected fingerprints (ensemble fingerprints) that summarize motif occurrences under the Boltzmann distribution $p_{S,beta}(G) = \frac{e^{-beta E(G)}}{Z_S(beta)}$. By applying these chemistry-informed fingerprints to SELEX data, the authors demonstrate robust community detection and subgraph-level explainability, enabling identification of low-abundance candidates that warrant experimental follow-up. The approach offers a principled framework for anomaly detection and rational candidate prioritization in aptamer discovery, while acknowledging limitations such as exclusion of pseudoknots and fixed thermodynamic parameters; future directions include learning parameters, exploring multi-temperature ensembles, and experimental validation.
Abstract
Machine-learning methods in biochemistry commonly represent molecules as graphs of pairwise intermolecular interactions for property and structure predictions. Most methods operate on a single graph, typically the minimal free energy (MFE) structure, for low-energy ensembles (conformations) representative of structures at thermodynamic equilibrium. We introduce a thermodynamically parameterized exponential-family random graph (ERGM) embedding that models molecules as Boltzmann-weighted ensembles of interaction graphs. We evaluate this embedding on SELEX datasets, where experimental biases (e.g., PCR amplification or sequencing noise) can obscure true aptamer-ligand affinity, producing anomalous candidates whose observed abundance diverges from their actual binding strength. We show that the proposed embedding enables robust community detection and subgraph-level explanations for aptamer ligand affinity, even in the presence of biased observations. This approach may be used to identify low-abundance aptamer candidates for further experimental evaluation.
