Table of Contents
Fetching ...

What do model reports say about their ChemBio benchmark evaluations? Comparing recent releases to the STREAM framework

Tom Reed, Tegan McCaslin, Luca Righetti

TL;DR

This paper investigates how three spring 2025 frontier AI model reports (OpenAI o3, Anthropic Claude Opus 4, and Google DeepMind Gemini 2.5 Pro) document ChemBio evaluations by mapping their content to STREAM v1 criteria. It finds substantial variability across reports, with 16 of 23 STREAM criteria at least partially satisfied by at least one report and eight fully satisfied in at least one instance, but also seven criteria unmet by any report. Case studies illustrate differences in human baselines, benchmark methodology variations, and private benchmark explanations, while highlighting gaps such as lack of example test items and detailed elicitation procedures. The study argues for greater transparency and cross-industry sharing of best practices to advance the science of AI evaluation, and suggests STREAM be used as a practical checklist to drive consistency and continuous improvement.

Abstract

Most frontier AI developers publicly document their safety evaluations of new AI models in model reports, including testing for chemical and biological (ChemBio) misuse risks. This practice provides a window into the methodology of these evaluations, helping to build public trust in AI systems, and enabling third party review in the still-emerging science of AI evaluation. But what aspects of evaluation methodology do developers currently include -- or omit -- in their reports? This paper examines three frontier AI model reports published in spring 2025 with among the most detailed documentation: OpenAI's o3, Anthropic's Claude 4, and Google DeepMind's Gemini 2.5 Pro. We compare these using the STREAM (v1) standard for reporting ChemBio benchmark evaluations. Each model report included some useful details that the others did not, and all model reports were found to have areas for development, suggesting that developers could benefit from adopting one another's best reporting practices. We identified several items where reporting was less well-developed across all model reports, such as providing examples of test material, and including a detailed list of elicitation conditions. Overall, we recommend that AI developers continue to strengthen the emerging science of evaluation by working towards greater transparency in areas where reporting currently remains limited.

What do model reports say about their ChemBio benchmark evaluations? Comparing recent releases to the STREAM framework

TL;DR

This paper investigates how three spring 2025 frontier AI model reports (OpenAI o3, Anthropic Claude Opus 4, and Google DeepMind Gemini 2.5 Pro) document ChemBio evaluations by mapping their content to STREAM v1 criteria. It finds substantial variability across reports, with 16 of 23 STREAM criteria at least partially satisfied by at least one report and eight fully satisfied in at least one instance, but also seven criteria unmet by any report. Case studies illustrate differences in human baselines, benchmark methodology variations, and private benchmark explanations, while highlighting gaps such as lack of example test items and detailed elicitation procedures. The study argues for greater transparency and cross-industry sharing of best practices to advance the science of AI evaluation, and suggests STREAM be used as a practical checklist to drive consistency and continuous improvement.

Abstract

Most frontier AI developers publicly document their safety evaluations of new AI models in model reports, including testing for chemical and biological (ChemBio) misuse risks. This practice provides a window into the methodology of these evaluations, helping to build public trust in AI systems, and enabling third party review in the still-emerging science of AI evaluation. But what aspects of evaluation methodology do developers currently include -- or omit -- in their reports? This paper examines three frontier AI model reports published in spring 2025 with among the most detailed documentation: OpenAI's o3, Anthropic's Claude 4, and Google DeepMind's Gemini 2.5 Pro. We compare these using the STREAM (v1) standard for reporting ChemBio benchmark evaluations. Each model report included some useful details that the others did not, and all model reports were found to have areas for development, suggesting that developers could benefit from adopting one another's best reporting practices. We identified several items where reporting was less well-developed across all model reports, such as providing examples of test material, and including a detailed list of elicitation conditions. Overall, we recommend that AI developers continue to strengthen the emerging science of evaluation by working towards greater transparency in areas where reporting currently remains limited.
Paper Structure (13 sections, 6 figures)

This paper contains 13 sections, 6 figures.

Figures (6)

  • Figure 1: STREAM criteria assessment for three model reports from spring 2025.
  • Figure 2: STREAM criteria assessment for three model reports from spring 2025. VCT* = VMQA, single select (SecureBio); LAB = LAB-Bench Subset - ProtocolQA, Cloning Scenarios, SeqQA (FutureHouse); WMD = Weapons of Mass Destruction Proxy, chem & bio datasets (Li et al., 2024); LAB* = ProtocolQA, open-ended (FutureHouse, OpenAI); LFB = Long-form Biorisk Questions (Gryphon Scientific, OpenAI); TTK = Tacit Knowledge & Troubleshooting (Gryphon Scientific, OpenAI); VCT = Virology Capabilities Test, multiple response (SecureBio); LAB† = LAB-Bench Subset - FigQA, ProtocolQA, Cloning Scenarios, SeqQA (FutureHouse); LFV = Long-form Virology Tasks (SecureBio, Deloitte, Signature Science, Anthropic); SSE = DNA Synthesis Screening Evasion (SecureBio); BKQ = Bioweapons knowledge questions (Deloitte); CrB = Creative Biology (SecureBio); SHB = Short-horizon computational biology tasks (Faculty.ai, Anthropic). Details of assessment and rationales can be found in Appendix 1.
  • Figure 3: Case study of human expert baselines. Panel A reproduces the ChemBio evaluation results figure from the Gemini 2.5 Pro model report. Sources: google_deepmind2025gemini Panel B modifies these to include human expert scores. Note that the human baseline scores for VMQA (VCT) were obtained for the harder ‘multiple-response’ version of this evaluation, while the model results are recorded in the ‘multiple choice’ mode. Sources for human expert scores: götting2025virologycapabilitiestestvctlaurent2024labbenchmeasuringcapabilitieslanguagedev2025comprehensive
  • Figure 4: Case study of variations in benchmark methodology. Panel A reproduces the Virology Capabilities Test results as presented by the o3 model report. Sources: openai2025o3 Panel B highlights how such scores change when switching from the 'single-select' to a ‘multiple-response’ version of this evaluation. Sources: götting2025virologycapabilitiestestvct
  • Figure 5: Summary of STREAM reporting requirements. Source: McCaslin et al. mccaslin2025streamchembiostandardtransparently
  • ...and 1 more figures