Table of Contents
Fetching ...

VERA-MH Concept Paper

Luca Belli, Kate Bentley, Will Alexander, Emily Ward, Matt Hawrilenko, Kelly Johnston, Mill Brown, Adam Chekroud

TL;DR

VERA-MH presents an automated, clinician-informed framework to evaluate the safety of AI chatbots in mental health, focusing initially on suicide risk. The system uses user-agent personas to simulate realistic conversations and a judge-agent to score interactions against a five-dimension rubric (Detects risk, Probes risk, Takes appropriate actions, Validates and collaborates, Maintains safe boundaries) within a $5\times4$ matrix, enabling scalable multi-turn assessments that are model-agnostic. Early results show GPT-5 achieving more Best Practice ratings than Claude variants, with ongoing human validation revealing strengths in validation/collaboration but variability in probing and clinician–judge alignment, guiding iterative improvements. Limitations include interpretability of outputs, potential score saturation, simulation fidelity, constrained persona diversity, and computational costs, with planned expansions to broader clinical validation and community feedback to enhance safety benchmarks for mental health AI tools.

Abstract

We introduce VERA-MH (Validation of Ethical and Responsible AI in Mental Health), an automated evaluation of the safety of AI chatbots used in mental health contexts, with an initial focus on suicide risk. Practicing clinicians and academic experts developed a rubric informed by best practices for suicide risk management for the evaluation. To fully automate the process, we used two ancillary AI agents. A user-agent model simulates users engaging in a mental health-based conversation with the chatbot under evaluation. The user-agent role-plays specific personas with pre-defined risk levels and other features. Simulated conversations are then passed to a judge-agent who scores them based on the rubric. The final evaluation of the chatbot being tested is obtained by aggregating the scoring of each conversation. VERA-MH is actively under development and undergoing rigorous validation by mental health clinicians to ensure user-agents realistically act as patients and that the judge-agent accurately scores the AI chatbot. To date we have conducted preliminary evaluation of GPT-5, Claude Opus and Claude Sonnet using initial versions of the VERA-MH rubric and used the findings for further design development. Next steps will include more robust clinical validation and iteration, as well as refining actionable scoring. We are seeking feedback from the community on both the technical and clinical aspects of our evaluation.

VERA-MH Concept Paper

TL;DR

VERA-MH presents an automated, clinician-informed framework to evaluate the safety of AI chatbots in mental health, focusing initially on suicide risk. The system uses user-agent personas to simulate realistic conversations and a judge-agent to score interactions against a five-dimension rubric (Detects risk, Probes risk, Takes appropriate actions, Validates and collaborates, Maintains safe boundaries) within a matrix, enabling scalable multi-turn assessments that are model-agnostic. Early results show GPT-5 achieving more Best Practice ratings than Claude variants, with ongoing human validation revealing strengths in validation/collaboration but variability in probing and clinician–judge alignment, guiding iterative improvements. Limitations include interpretability of outputs, potential score saturation, simulation fidelity, constrained persona diversity, and computational costs, with planned expansions to broader clinical validation and community feedback to enhance safety benchmarks for mental health AI tools.

Abstract

We introduce VERA-MH (Validation of Ethical and Responsible AI in Mental Health), an automated evaluation of the safety of AI chatbots used in mental health contexts, with an initial focus on suicide risk. Practicing clinicians and academic experts developed a rubric informed by best practices for suicide risk management for the evaluation. To fully automate the process, we used two ancillary AI agents. A user-agent model simulates users engaging in a mental health-based conversation with the chatbot under evaluation. The user-agent role-plays specific personas with pre-defined risk levels and other features. Simulated conversations are then passed to a judge-agent who scores them based on the rubric. The final evaluation of the chatbot being tested is obtained by aggregating the scoring of each conversation. VERA-MH is actively under development and undergoing rigorous validation by mental health clinicians to ensure user-agents realistically act as patients and that the judge-agent accurately scores the AI chatbot. To date we have conducted preliminary evaluation of GPT-5, Claude Opus and Claude Sonnet using initial versions of the VERA-MH rubric and used the findings for further design development. Next steps will include more robust clinical validation and iteration, as well as refining actionable scoring. We are seeking feedback from the community on both the technical and clinical aspects of our evaluation.
Paper Structure (17 sections, 6 figures, 3 tables)

This paper contains 17 sections, 6 figures, 3 tables.

Figures (6)

  • Figure 1: VERA-MH overall design.
  • Figure 2: Evaluation of Claude Sonnet as provider.
  • Figure 3: Evaluation of Claude Opus as provider.
  • Figure 4: Evaluation of ChatGPT-5 as provider.
  • Figure 5: Percent of rated conversations (N = 75) with matched clinician and judge-agent ratings, by dimension.
  • ...and 1 more figures