Table of Contents
Fetching ...

LacMaterial: Large Language Models as Analogical Chemists for Materials Discovery

Hongyu Guo

TL;DR

The paper investigates whether large language models can perform explicit cross-domain and in-domain analogical reasoning to accelerate materials discovery, focusing on solid-state Li-conducting electrolytes. It introduces cross-domain analogy prompting, analogy-guided exemplars, and majority-vote selection to generate novel LLZO-like candidates beyond conventional substitutions, and compares with naive prompts. It also builds in-domain templates from a larger labeled corpus to guide exploration within the Li-conductor domain, demonstrating interpolations between known compositions. Computational validations using surrogate energies indicate thermodynamic viability of candidate designs, while acknowledging the challenge of computing ionic conductivities. The work highlights a path toward interpretable, analogy-driven hypothesis generation and potential closed-loop discovery with surrogate models.

Abstract

Analogical reasoning, the transfer of relational structures across contexts (e.g., planet is to sun as electron is to nucleus), is fundamental to scientific discovery. Yet human insight is often constrained by domain expertise and surface-level biases, limiting access to deeper, structure-driven analogies both within and across disciplines. Large language models (LLMs), trained on vast cross-domain data, present a promising yet underexplored tool for analogical reasoning in science. Here, we demonstrate that LLMs can generate novel battery materials by (1) retrieving cross-domain analogs and analogy-guided exemplars to steer exploration beyond conventional dopant substitutions, and (2) constructing in-domain analogical templates from few labeled examples to guide targeted exploitation. These explicit analogical reasoning strategies yield candidates outside established compositional spaces and outperform standard prompting baselines. Our findings position LLMs as interpretable, expert-like hypothesis generators that leverage analogy-driven generalization for scientific innovation.

LacMaterial: Large Language Models as Analogical Chemists for Materials Discovery

TL;DR

The paper investigates whether large language models can perform explicit cross-domain and in-domain analogical reasoning to accelerate materials discovery, focusing on solid-state Li-conducting electrolytes. It introduces cross-domain analogy prompting, analogy-guided exemplars, and majority-vote selection to generate novel LLZO-like candidates beyond conventional substitutions, and compares with naive prompts. It also builds in-domain templates from a larger labeled corpus to guide exploration within the Li-conductor domain, demonstrating interpolations between known compositions. Computational validations using surrogate energies indicate thermodynamic viability of candidate designs, while acknowledging the challenge of computing ionic conductivities. The work highlights a path toward interpretable, analogy-driven hypothesis generation and potential closed-loop discovery with surrogate models.

Abstract

Analogical reasoning, the transfer of relational structures across contexts (e.g., planet is to sun as electron is to nucleus), is fundamental to scientific discovery. Yet human insight is often constrained by domain expertise and surface-level biases, limiting access to deeper, structure-driven analogies both within and across disciplines. Large language models (LLMs), trained on vast cross-domain data, present a promising yet underexplored tool for analogical reasoning in science. Here, we demonstrate that LLMs can generate novel battery materials by (1) retrieving cross-domain analogs and analogy-guided exemplars to steer exploration beyond conventional dopant substitutions, and (2) constructing in-domain analogical templates from few labeled examples to guide targeted exploitation. These explicit analogical reasoning strategies yield candidates outside established compositional spaces and outperform standard prompting baselines. Our findings position LLMs as interpretable, expert-like hypothesis generators that leverage analogy-driven generalization for scientific innovation.
Paper Structure (12 sections, 1 figure, 7 tables)