Table of Contents
Fetching ...

An Encoder-Decoder Foundation Chemical Language Model for Generative Polymer Design

Harikrishna Sahu, Wei Xiong, Anagha Savit, Shivank S Shukla, Rampi Ramprasad

TL;DR

PolyT5, an encoder-decoder chemical language model based on the T5 architecture, trained to understand and generate polymer structures, is presented, enabling both property prediction and the targeted generation of polymers conditioned on desired property values.

Abstract

Traditional machine learning has advanced polymer discovery, yet direct generation of chemically valid and synthesizable polymers without exhaustive enumeration remains a challenge. Here we present polyT5, an encoder-decoder chemical language model based on the T5 architecture, trained to understand and generate polymer structures. polyT5 enables both property prediction and the targeted generation of polymers conditioned on desired property values. We demonstrate its utility for dielectric polymer design, seeking candidates with dielectric constant >3, bandgap >4 eV, and glass transition temperature >400 K, alongside melt-processability and solubility requirements. From over 20,000 generated promising candidates, one was experimentally synthesized and validated, showing strong agreement with predictions. To further enhance usability, we integrated polyT5 within an agentic AI framework that couples it with a general-purpose LLM, allowing natural language interaction for property prediction and generative design. Together, these advances establish a versatile and accessible framework for accelerated polymer discovery.

An Encoder-Decoder Foundation Chemical Language Model for Generative Polymer Design

TL;DR

PolyT5, an encoder-decoder chemical language model based on the T5 architecture, trained to understand and generate polymer structures, is presented, enabling both property prediction and the targeted generation of polymers conditioned on desired property values.

Abstract

Traditional machine learning has advanced polymer discovery, yet direct generation of chemically valid and synthesizable polymers without exhaustive enumeration remains a challenge. Here we present polyT5, an encoder-decoder chemical language model based on the T5 architecture, trained to understand and generate polymer structures. polyT5 enables both property prediction and the targeted generation of polymers conditioned on desired property values. We demonstrate its utility for dielectric polymer design, seeking candidates with dielectric constant >3, bandgap >4 eV, and glass transition temperature >400 K, alongside melt-processability and solubility requirements. From over 20,000 generated promising candidates, one was experimentally synthesized and validated, showing strong agreement with predictions. To further enhance usability, we integrated polyT5 within an agentic AI framework that couples it with a general-purpose LLM, allowing natural language interaction for property prediction and generative design. Together, these advances establish a versatile and accessible framework for accelerated polymer discovery.
Paper Structure (24 sections, 18 figures, 6 tables)

This paper contains 24 sections, 18 figures, 6 tables.

Figures (18)

  • Figure 1: Schematic workflow illustrating the large language model (LLM)-based framework for organic material design targeting dielectric applications. (A) polyT5, a T5-based language model, pre-trained on a corpus of 100 million polymer structures to capture underlying chemical and structural patterns. (B) Fine-tuning polyT5 for domain-specific tasks, including: (i) property prediction for thermal, electrical, and solubility-related properties and (ii) generative design — generating hypothetical polymer candidates based on a target property (e.g., glass transition temperature, $T_{\rm g}$)). (C) Application of the fine-tuned model for dielectric polymer design by generating hypothetical polymers targeting a $T_{\rm g}$ of 500 K, followed by screening to identify candidates satisfying dielectric, thermal stability, and solubility criteria, thus enabling a data-driven materials discovery pipeline.
  • Figure 2: polyT5: Key steps in the design of polymers for dielectric applications. (A) Conversion of polymer SMILES (PSMILES) to polymer SELFIES (PSELFIES) representations. (B) Tokenization of PSELFIES for language model input. (C) Implementation of a masking strategy for pre-training the base polyT5 model. (D) Fine-tuning of the base model for property prediction tasks, including thermal, electronic, and solubility properties. (E) Fine-tuning of the base model for conditional hypothetical polymer generation. (F) Generation and screening of hypothetical polymers targeting desired dielectric properties through integrated property prediction models.
  • Figure 3: Performance of fine-tuned polyT5-medium models for various property prediction tasks. (A) Learning curves showing the root-mean-square errors (RMSEs) for thermal and electronic property predictions. (B) Learning curves for solubility prediction. For panels A and B, error bars represent the standard deviation from five different random train-test splits. (C--G) Parity plots for a representative split, comparing predicted and experimental values for glass transition temperature ($T_{\rm g}$), thermal decomposition temperature ($T_d$), melting temperature ($T_m$), electronic band gap ($E_g$), and dielectric constant ($\varepsilon$). (H) Confusion matrix for a representative split for solubility prediction, with values reported in percentage.
  • Figure 4: Hypothetical candidate generation using polyT5. (A) Optimal combination of fine-tuning epoch, T5 sampling temperature, and sampling probability identified for small (14, 0.9, 0.75), medium (6, 1.1, 0.75), and large (8, 1.1, 0.95) polyT5 models for maximizing the generation of valid hypothetical polymers. SV, TSD, DD, and PV represent successive validation filters: validity of SMILES strings verified by RDKit (SV), removal of duplicates against the training set (TSD), deduplication within the generated hypothetical polymer dataset (DD), and polymer validity (PV) ensuring the presence of exactly two Astatine (At) atoms each with valency one. Note that PV $\subset$ DD $\subset$ TSD $\subset$ SV. (B) Distribution of Tanimoto similarity values between each of the over 6 million polyT5-generated hypothetical polymers and their closest known polymer in the training dataset, as a function of the closest polymer’s $T_{\rm g}$ (K). (C) Distribution of polymer classes based on the presence of functional groups in the 6 million generated hypothetical polymers. (D) Distribution of predicted $T_\mathrm{g}$ values by polyT5-medium for over 6 million candidate polymers.
  • Figure 5: Designing dielectric polymers with polyT5. (A) Screening of candidates based on thermal, electronic, and solubility criteria. (B) Synthetic accessibility scores of selected polymers versus predicted $T_\mathrm{g}$, annotated by SELFIES reproducibility, with corresponding marginal histograms. (C) Selected polymer for experimental validation
  • ...and 13 more figures