Table of Contents
Fetching ...

Protein generation with embedding learning for motif diversification

Kevin Michalewicz, Chen Jin, Philip Alexander Teare, Tom Diethe, Mauricio Barahona, Barbara Bravi, Asher Mullokandov

TL;DR

PGEL addresses the diversity–fidelity trade-off in motif-focused protein design by learning a high-dimensional motif embedding $v_*$ within a frozen diffusion model’s denoiser and perturbing it in embedding space to diversify motifs while preserving scaffold geometry. By combining embedding masking of scaffold MSA features with a loss that enforces backbone fidelity and plausible torsion, PGEL achieves greater designability and structural diversity than partial diffusion across three representative cases, with self-consistency demonstrated via inverse folding and AlphaFold3 in several designs. Although results are in silico, the approach provides a general, scalable strategy to systematically diversify functional motifs without retraining diffusion models, potentially accelerating motif-engineering workflows in protein design. The study highlights embedding-centric diversification as a practical alternative to geometry-centric perturbations, offering broader applicability to pre-trained diffusion models for proteins.

Abstract

A fundamental challenge in protein design is the trade-off between generating structural diversity while preserving motif biological function. Current state-of-the-art methods, such as partial diffusion in RFdiffusion, often fail to resolve this trade-off: small perturbations yield motifs nearly identical to the native structure, whereas larger perturbations violate the geometric constraints necessary for biological function. We introduce Protein Generation with Embedding Learning (PGEL), a general framework that learns high-dimensional embeddings encoding sequence and structural features of a target motif in the representation space of a diffusion model's frozen denoiser, and then enhances motif diversity by introducing controlled perturbations in the embedding space. PGEL is thus able to loosen geometric constraints while satisfying typical design metrics, leading to more diverse yet viable structures. We demonstrate PGEL on three representative cases: a monomer, a protein-protein interface, and a cancer-related transcription factor complex. In all cases, PGEL achieves greater structural diversity, better designability, and improved self-consistency, as compared to partial diffusion. Our results establish PGEL as a general strategy for embedding-driven protein generation allowing for systematic, viable diversification of functional motifs.

Protein generation with embedding learning for motif diversification

TL;DR

PGEL addresses the diversity–fidelity trade-off in motif-focused protein design by learning a high-dimensional motif embedding within a frozen diffusion model’s denoiser and perturbing it in embedding space to diversify motifs while preserving scaffold geometry. By combining embedding masking of scaffold MSA features with a loss that enforces backbone fidelity and plausible torsion, PGEL achieves greater designability and structural diversity than partial diffusion across three representative cases, with self-consistency demonstrated via inverse folding and AlphaFold3 in several designs. Although results are in silico, the approach provides a general, scalable strategy to systematically diversify functional motifs without retraining diffusion models, potentially accelerating motif-engineering workflows in protein design. The study highlights embedding-centric diversification as a practical alternative to geometry-centric perturbations, offering broader applicability to pre-trained diffusion models for proteins.

Abstract

A fundamental challenge in protein design is the trade-off between generating structural diversity while preserving motif biological function. Current state-of-the-art methods, such as partial diffusion in RFdiffusion, often fail to resolve this trade-off: small perturbations yield motifs nearly identical to the native structure, whereas larger perturbations violate the geometric constraints necessary for biological function. We introduce Protein Generation with Embedding Learning (PGEL), a general framework that learns high-dimensional embeddings encoding sequence and structural features of a target motif in the representation space of a diffusion model's frozen denoiser, and then enhances motif diversity by introducing controlled perturbations in the embedding space. PGEL is thus able to loosen geometric constraints while satisfying typical design metrics, leading to more diverse yet viable structures. We demonstrate PGEL on three representative cases: a monomer, a protein-protein interface, and a cancer-related transcription factor complex. In all cases, PGEL achieves greater structural diversity, better designability, and improved self-consistency, as compared to partial diffusion. Our results establish PGEL as a general strategy for embedding-driven protein generation allowing for systematic, viable diversification of functional motifs.
Paper Structure (15 sections, 5 equations, 5 figures, 3 tables, 3 algorithms)

This paper contains 15 sections, 5 equations, 5 figures, 3 tables, 3 algorithms.

Figures (5)

  • Figure 1: Outline of the PGEL learning and generation procedures during one reverse diffusion step.
  • Figure 2: Summary of the evaluation metrics. PGEL takes as input a PDB entry and the amino acid range corresponding to the motif. From 1000 PGEL-generated backbones, designable candidates are filtered by root mean square deviation (RMSD) and pLDDT thresholds, and structural diversity is assessed via hierarchical clustering. Backbones are also required to be distinguishable from the native. Cluster representatives undergo self-consistency evaluation: sequences assigned to the designable backbones with ProteinMPNN are refolded, and at least one predicted structure must satisfy set mRMSD and predicted alignment error (pAE) conditions relative to the generated backbone.
  • Figure 3: Results for example 1. (A) Number of clusters, both total and distinguishable from native, as a function of the TM-score threshold $t_m$ for PGEL and partial diffusion. (B) Left: two successful PGEL designs at $t_m=0.6$ using column masking with rates $\alpha=0.5$ and $\alpha=0.75$. Right: two partial diffusion failed backbones at $t_m=0.6$, one obtained with $T=25$ timesteps that violates the motif RMSD constraint, and one with $T=2$ timesteps that is not distinguishable from the native.
  • Figure 4: (A) Example 2: Left: Number of clusters identified PGEL and partial diffusion in the binding site example, both total and distinguishable from native, as a function of the TM-score threshold $t_m$. Right: generated binding site backbones (overlaid with native), alongside the native sequence and sequences that refold to self-consistent structures. (B) Same as A but for example 3.
  • Figure 5: Ratios of the binding affinity predicted with PRODIGY of the native structure vs the AlphaFold3-predicted structures from backbones (one per cluster) generated by PGEL and partial diffusion for example 2 (A) and example 3 (B).