FAQ

Frequently Asked Questions

Why was VERA-MH developed?

People are turning to AI for mental health support. Without clear safeguards, some AI tools can increase distress, reinforce harmful thoughts, and miss warning signals. As cases of real-world harm emerged, it became clear that the field needed collaboratively developed, clinically grounded, safety standards to reliably protect people in their most vulnerable moments.

This urgent unmet need led to the creation of VERA-MH. Open-source safety standards help measure risk, identify gaps, and improve how AI tools respond when someone may be at risk.

Spring Health worked in close collaboration with the AI in Mental Health Safety & Ethics Council, a coalition of experts, to create the initial standards, which were then improved with feedback from AI experts, technologists, clinicians, and organizations who share a similar commitment to AI safety.

How does VERA-MH  evaluate safety?

VERA-MH works in two steps by simulating multiple AI conversations with different individuals experiencing different levels of suicide risk.

First, a “user agent” (an AI model) plays the role of a person using one of many realistic profiles (background, mental health conditions, demographics, and communication styles). The AI tool responds to input in real time.

Next, a separate “judge agent” reviews the resulting multi-turn conversation and scores the AI tool against the rubric. The rubric is a clinically validated score card, informed by high safety standards and suicide prevention best practices.

The scoring rubric is built on best-practice clinical guidance and designed so that different real-life expert human clinicians would score the same conversation in the same way. VERA-MH applies those same rules to its judge agent, producing consistent, dependable scores you can trust when comparing one AI tool to another.

What does VERA-MH measure?

The VERA-MH tool scores an AI tool on how well it:

  • Detects Potential Risk: Does the tool detect statements indicating the user is at potential risk of suicide?
  • Confirms Risk: Does the tool ask follow-up questions when needed to determine whether the individual is having suicidal thoughts?
  • Guides to Human Care: Does the tool provide appropriate resources and guide to human support when risk is identified?
  • Holds a Supportive Conversation: Does the tool use an appropriate tone, style of communication, and level of validation?
  • Follows AI Boundaries: Does the tool remind of the limitations of AI and avoid fueling potentially harmful behavior?

Who can use VERA-MH?

Our code is open-source, so any developer or researcher can plug VERA-MH code into their AI tool to receive a safety score and easily determine how well and safely a tool responds to conversations involving suicide risk.

  • Developers can use VERA-MH to get better guidance on what safe AI looks like, helping them spot problems and make improvements faster.
  • Employers and health plans should require VERA-MH scores to establish a consistent, clinical benchmark for AI safety. This gives buyers a consistent, clinical benchmark for vendor oversight and helps manage risk as AI adoption scales.
  • Benefits consultants can more consistently and fairly evaluate AI mental health solutions and make informed suggestions by requesting VERA-MH scores as part of client RFPs.
  • Researchers and Policymakers gain a common language to create guidelines, oversight, and future regulations.

Why is VERA-MH a clinically grounded standard for AI safety in mental health?

  • VERA-MH applies more rigorous, clinically grounded safety benchmarks than other evaluation tools available today.
  • VERA-MH's safety determinations align with those of licensed clinicians who rated the same conversations using the same scoring rubric.
  • VERA-MH has been developed in partnership with many external, objective stakeholders (clinicians, developers, vendors, suicide prevention and mental health experts).
  • The AI in Mental Health Safety & Ethics Council and Spring Health researchers sought and incorporated input from a broad range of external experts during a request for feedback period.
  • VERA-MH is entirely open-source and automated which allows for ongoing evaluation criteria updating as guidelines and clinical best practices evolve.

How does VERA-MH compare to expert human clinician scoring?

Research shows that the VERA-MH AI judge scoring conversations consistently aligns with the judgment of expert clinicians. In this study, the AI matched independent clinician scoring, performing at a level of reliability comparable to the human "gold standard."

What’s next for VERA-MH?

Throughout 2026, the VERA-MH team will continue to publish peer-reviewed papers as it expands the benchmark beyond suicide risk to additional safety domains, including ongoing work on harm to others. The team also plans to begin using real-world data to further validate VERA-MH.

How can I get involved with VERA-MH as a developer?

There are several meaningful ways to participate:

  1. Run VERA-MH on your own AI tools: Download the open-source VERA-MH code and run the evaluation on your AI tools. This provides a standard rating for high-risk mental health scenarios and identifies areas for safety improvement.
  2. Share feedback and help shape what’s next: VERA-MH is designed to evolve with the community. Submit feedback through this link to help refine the framework.
  3. Contribute to the development of the code: Submit contributions to the github repository.
  4. Share results: Post your VERA-MH scores. Transparency helps the community learn together and move toward making safety a real, shared standard, not just a claim.

What questions should I ask when assessing the safety of AI as an employer or as a benefits consultant?

Use the following questions in RFIs and RFPs to better understand the AI safety and security of vendor products:

  • Is there a 24/7/365 defined human clinician escalation path for ambiguous or high-risk cases?
  • Do you have a multi-layer AI safety framework?
  • Do you have a zero-retention policy to ensure AI systems donʼt store or use data for training purposes?
  • What governance, compliance, and transparency controls are in place?
  • Is the AI assisting clinicians or replacing clinical judgment?
  • Are members explicitly informed when they are interacting with AI, how it’s being used, and whether they can choose a human-only interaction?
  • What independent evidence demonstrates that the AI is safe, especially in high-risk cases?
  • How are models monitored, updated, and governed over time?
  • What is the VERA-MH safety score for the mental health tool?

What does it mean that VERA-MH is clinically validated?

For VERA-MH, "clinically validated” means two things right now. First, practicing clinicians built the scoring rubric, with input from external clinicians and suicide prevention specialists. Second, we tested it: clinicians and VERA-MH's AI judge scored the same simulated conversations against that rubric, and the judge's ratings closely matched the clinicians'.

Two limitations worth being clear about. The current version of VERA-MH covers only suicide risk, not mental health safety more broadly. Validation so far is based on simulated conversations, not real ones.

Looking ahead, we plan to validate real-world conversations and link scores to actual user outcomes, which will be the strongest test of whether a higher score really means someone is safer.