Table of Contents
Fetching ...

Generative AI for Biosciences: Emerging Threats and Roadmap to Biosecurity

Zaixi Zhang, Souradip Chakraborty, Amrit Singh Bedi, Emilin Mathew, Varsha Saravanan, Le Cong, Alvaro Velasquez, Sheng Lin-Gibson, Megan Blewett, Dan Hendrycs, Alex John London, Ellen Zhong, Ben Raphael, Adji Bousso Dieng, Jian Ma, Eric Xing, Russ Altman, George Church, Mengdi Wang

TL;DR

The paper addresses the dual-use risks that GenAI introduces to biosciences, arguing that existing safeguards lag behind capabilities like autonomous agents and open-ended design. It synthesizes insights from 130 expert interviews and surveys to categorize threats (jailbreaks, privacy leaks, autonomous agents) and maps a lifecycle-based safety roadmap spanning pre-training, post-training, and inference, plus adaptive governance. Key contributions include a taxonomy of frontier risks, evidence on gaps in current screening and governance, and concrete multi-layer defenses (data curation, watermarking, RLHF/adversarial training, anti-jailbreak screening, red-teaming, and inference-time alignment) along with calls for global coordination and secure-by-design practices. The practical impact is a actionable blueprint for researchers, policymakers, and industry to secure GenAI-enabled bioscience while preserving acceleration of discovery and innovation.

Abstract

The rapid adoption of generative artificial intelligence (GenAI) in the biosciences is transforming biotechnology, medicine, and synthetic biology. Yet this advancement is intrinsically linked to new vulnerabilities, as GenAI lowers the barrier to misuse and introduces novel biosecurity threats, such as generating synthetic viral proteins or toxins. These dual-use risks are often overlooked, as existing safety guardrails remain fragile and can be circumvented through deceptive prompts or jailbreak techniques. In this Perspective, we first outline the current state of GenAI in the biosciences and emerging threat vectors ranging from jailbreak attacks and privacy risks to the dual-use challenges posed by autonomous AI agents. We then examine urgent gaps in regulation and oversight, drawing on insights from 130 expert interviews across academia, government, industry, and policy. A large majority ($\approx 76$\%) expressed concern over AI misuse in biology, and 74\% called for the development of new governance frameworks. Finally, we explore technical pathways to mitigation, advocating a multi-layered approach to GenAI safety. These defenses include rigorous data filtering, alignment with ethical principles during development, and real-time monitoring to block harmful requests. Together, these strategies provide a blueprint for embedding security throughout the GenAI lifecycle. As GenAI becomes integrated into the biosciences, safeguarding this frontier requires an immediate commitment to both adaptive governance and secure-by-design technologies.

Generative AI for Biosciences: Emerging Threats and Roadmap to Biosecurity

TL;DR

The paper addresses the dual-use risks that GenAI introduces to biosciences, arguing that existing safeguards lag behind capabilities like autonomous agents and open-ended design. It synthesizes insights from 130 expert interviews and surveys to categorize threats (jailbreaks, privacy leaks, autonomous agents) and maps a lifecycle-based safety roadmap spanning pre-training, post-training, and inference, plus adaptive governance. Key contributions include a taxonomy of frontier risks, evidence on gaps in current screening and governance, and concrete multi-layer defenses (data curation, watermarking, RLHF/adversarial training, anti-jailbreak screening, red-teaming, and inference-time alignment) along with calls for global coordination and secure-by-design practices. The practical impact is a actionable blueprint for researchers, policymakers, and industry to secure GenAI-enabled bioscience while preserving acceleration of discovery and innovation.

Abstract

The rapid adoption of generative artificial intelligence (GenAI) in the biosciences is transforming biotechnology, medicine, and synthetic biology. Yet this advancement is intrinsically linked to new vulnerabilities, as GenAI lowers the barrier to misuse and introduces novel biosecurity threats, such as generating synthetic viral proteins or toxins. These dual-use risks are often overlooked, as existing safety guardrails remain fragile and can be circumvented through deceptive prompts or jailbreak techniques. In this Perspective, we first outline the current state of GenAI in the biosciences and emerging threat vectors ranging from jailbreak attacks and privacy risks to the dual-use challenges posed by autonomous AI agents. We then examine urgent gaps in regulation and oversight, drawing on insights from 130 expert interviews across academia, government, industry, and policy. A large majority (\%) expressed concern over AI misuse in biology, and 74\% called for the development of new governance frameworks. Finally, we explore technical pathways to mitigation, advocating a multi-layered approach to GenAI safety. These defenses include rigorous data filtering, alignment with ethical principles during development, and real-time monitoring to block harmful requests. Together, these strategies provide a blueprint for embedding security throughout the GenAI lifecycle. As GenAI becomes integrated into the biosciences, safeguarding this frontier requires an immediate commitment to both adaptive governance and secure-by-design technologies.
Paper Structure (17 sections, 3 figures, 3 tables)

This paper contains 17 sections, 3 figures, 3 tables.

Figures (3)

  • Figure 1: Timeline of Emerging Generative AI (GenAI) for Biosciences (Pre-2020 to 2025). Squares denote GenAI models for nucleic acids (DNA and RNA); triangles represent GenAI models for proteins (sequences, structures, functions, etc.); circles indicate GenAI models for omics data (e.g., single-cell RNA-seq); pentagons mark GenAI models for small molecules; and diamonds signify foundation models for medical data (e.g., pathological images, electronic health records (EHR), and clinical trial data).
  • Figure 2: Emerging biosecurity threats of GenAI.(a) Structure-based protein design tools (e.g., RFDiffusion, AlphaFold) can be repurposed to engineer toxic proteins or viral components. (b) Genome foundation models such as DNABert and Evo could facilitate genetic modification of viral genomes, enhancing virulence or enabling immune escape. (c) Omics foundation models, including GeneFormer and scGPT, carry risks of reconstructing sensitive genetic or health-related information via privacy attacks. (d) Small-molecule generators like SyntheMol have the potential to design novel toxic compounds. (e) Medical foundation models may leak protected patient data from their training sets through membership or property inference attacks. (f) AI-driven scientist/agent platforms (e.g., ChemCrow, Biomni, OriGene) may autonomously accelerate threat design.
  • Figure 3: Overview of expert perspectives on the intersection of AI and biosecurity. (a) Distribution of 130 interviewees across four key sectors: Industry (n=43), Academia (n=41), Government (n=29), and Policy (n=17), with examples of representative institutions. (b) Methodology overview, outlining the semi-structured interview protocol with guiding questions and the primary themes of analysis derived from the responses. (c) Sector-specific analysis showing the percentage of interviewees within each sector who affirmed four key propositions: the urgency of AI misuse in biology, the inadequacy of current screening protocols, the need for governance, and support for functional screening. (d) Co-occurrence matrix of key themes across all 130 interviewees. The values indicate the number of individuals who hold both intersecting views.