Synergizing chemical and AI communities for advancing laboratories of the future
Saejin Oh, Xinyi Fang, I-Hsin Lin, Paris Dee, Christopher S. Dunham, Stacy M. Copp, Abigail G. Doyle, Javier Read de Alaniz, Mengyang Gu
TL;DR
This paper addresses accelerating chemical discovery by learning unknown relationships from digitized experimental data using ML and AI within automated laboratories. It proposes an integrated workflow that combines automated data collection, diverse predictive models, Bayesian optimization, and LLM agents to facilitate cross-disciplinary collaboration, framed around mappings such as $f(\mathbf x)$ and optimization objectives like $\mathbf x^* = \arg\max_{\mathbf x} g(\mathbf x)$. The authors present three case studies—physics-informed ML for block copolymer SAXS phase ID, ML-guided design of DNA-stabilized silver nanoclusters, and open-source Bayesian optimization for reaction development—to demonstrate speedups, improved design success, and practical workflows. The outlook highlights the transformative potential of self-driving labs while underscoring challenges in data access, standardization, and AI reliability that must be addressed to fully realize these systems in chemistry.
Abstract
The development of automated experimental facilities and the digitization of experimental data have introduced numerous opportunities to radically advance chemical laboratories. As many laboratory tasks involve predicting and understanding previously unknown chemical relationships, machine learning (ML) approaches trained on experimental data can substantially accelerate the conventional design-build-test-learn process. This outlook article aims to help chemists understand and begin to adopt ML predictive models for a variety of laboratory tasks, including experimental design, synthesis optimization, and materials characterization. Furthermore, this article introduces how artificial intelligence (AI) agents based on large language models can help researchers acquire background knowledge in chemical or data science and accelerate various aspects of the discovery process. We present three case studies in distinct areas to illustrate how ML models and AI agents can be leveraged to reduce time-consuming experiments and manual data analysis. Finally, we highlight existing challenges that require continued synergistic effort from both experimental and computational communities to address.
