Stanford Study Exposes Major Flaw in AI Mental Health Safety Testing | Stanford HAI
Stanford
University
  • Stanford Home
  • Maps & Directions
  • Search Stanford
  • Emergency Info
  • Terms of Use
  • Privacy
  • Copyright
  • Trademarks
  • Non-Discrimination
  • Accessibility
© Stanford University.  Stanford, California 94305.
Skip to content
  • About

    • About
    • People
    • Get Involved with HAI
    • Support HAI
    • Subscribe to Email
  • Research

    • Research
    • Fellowship Programs
    • Grants
    • Student Affinity Groups
    • Centers & Labs
    • Research Publications
    • Research Partners
  • Education

    • Education
    • Executive and Professional Education
    • Government and Policymakers
    • K-12
    • Stanford Students
  • Policy

    • Policy
    • Policy Publications
    • Policymaker Education
    • Student Opportunities
  • AI Index

    • AI Index
    • AI Index Report
    • Global Vibrancy Tool
    • People
  • News
  • Events
  • Industry
  • Centers & Labs

Stay Up To Date

Get the latest news, advances in research, policy work, and education program updates from HAI in your inbox weekly.

Sign Up For Latest News

Navigate
  • About
  • Events
  • AI Glossary
  • Careers
  • Search
Participate
  • Get Involved
  • Support HAI
  • Contact Us
news

Stanford Study Exposes Major Flaw in AI Mental Health Safety Testing

Date
July 13, 2026
Topics
Healthcare
Generative AI
Privacy, Safety, Security
mental health ai illustration head with binary code

With increased use of chatbots in mental health contexts, AI developers now rely on human experts to evaluate AI’s responses for “safety” – but experts rarely agree on what’s safe.

Faced with the fact that many AI users treat their chatbots as life coaches, counselors, and therapists, developers of the biggest and best-known large language models now employ psychologists and psychiatrists as “safety experts” to guide the training of new models in these nuanced and high-risk contexts. 

Currently, developers subject each new model to a battery of benchmark mental health queries. Human experts rate the chatbot’s responses on a quantitative scale, grading its relative success in providing “safe” advice to questions. In theory, AI developers can then use the ratings to hone the model’s performance to be safer.

But what if the experts disagree? 

This was the question Stanford researchers posed in a recent study of AI safety approaches in the mental health space. They asked three board-certified psychiatrists to evaluate 360 AI responses to synthetic mental health-related user prompts that the authors created for the study. None contained real user data or any personally identifiable information to protect privacy. Too often, however, the experts differed in their ratings, leaving AI developers struggling to find ways to improve AI mental health safety.

“The need for AI to get things right is especially great in the mental health space, and the way developers test their models for safety poses a danger for users. And the problem is worst in the areas of highest risk – when the users are suicidal or are in danger of self-harm,” says Kiana Jafari, a postdoctoral scholar at Stanford, director of the Stanford Center for AI Safety, and first author of the study. The paper, partially supported by the Stanford Institute for Human-Centered AI, was accepted to ACM FAccT 2026 and was presented at the American Psychiatric Association (APA) Annual Meeting 2026.  

Reversion to the Mean

Many assume that simply averaging the experts’ scores would provide an adequate baseline, but the researchers found something altogether different – averaging the scores only complicated matters, providing answers that were no one’s idea of a good response. 

“It doesn’t matter how many experts you have – 3, 10, or 1,000 – when they do not agree, you are not actually getting to the ground truth by averaging their scores,” Jafari says. “You end up steering your model toward no one’s ideal at all.”

“The disagreement is structural, not just noise or even bias in the data,” adds Nina Vasan, clinical assistant professor of psychiatry and behavioral sciences at Stanford School of Medicine and a co-author of the study. “You can add more experts, but we can’t seem to bridge the gap mathematically when the experts disagree.”

At a recent presentation of their findings at the annual meeting of the American Psychiatric Association, the researchers polled over 100 psychiatrists in attendance, and even with this much higher number of experts, the results were the same. Responses were almost evenly split across the board on their ratings of safety, empathy, and correctness.

None of the experts is necessarily wrong in their disparate opinions, Vasan explains. In fact, they are each right in their own ways, applying equally valid professional rubrics in their evaluations. “It’s a matter of professional judgment,” Vasan says.

To confirm this finding, the researchers interviewed their experts after the fact to find that disagreements typically stem from the experts’ clinical training and the frameworks they apply to diagnosis, treatment, and management. And, when the experts cannot agree, rating reliability falls below acceptable safety thresholds. 

“In high-risk mental health applications, such as users who are struggling with suicidal thoughts, psychosis, or eating disorders, AI safety is not yet there, and the developers must find new ways to train their models,” Vasan says.

Remedies

Ultimately, the researchers are spotlighting the concern for AI developers and mental health providers both, urging greater attention to AI safety and encouraging collaboration across fields to address these concerns. 

In the paper, they offer several interrelated remedies. The first is to demand greater transparency from AI developers about their reliability metrics and to declare which specific frameworks were used to test new models. 

Second, as the three clinical frameworks in use today – safety-first, engagement-centered, and culturally informed orientations – are incompatible and can’t be averaged, AI developers should model each framework individually and use it contextually where it is most appropriate.

Third, the field should see expert disagreement as a reflection of the complexity of the challenge and use it as a red flag to escalate discrepancies for greater human attention. 

“Preserve the disagreement. Don’t average it away,” Jafari says. “Expert disagreement isn’t a measurement problem. It’s a fundamental reality we need to understand and incorporate into system design before it’s a risk to users’ mental health.”

Share
Link copied to clipboard!
Contributor(s)
Andrew Myers

Related News

The Complexities of Governing Mental Health AI
Caroline Yee, Caroline Meinhardt, Michelle Mello, Jane Paik Kim
Jul 24, 2026
News
digital face mental health illustration

Policymakers, academics, healthcare providers, AI developers, and patient advocates convened by Stanford HAI identify critical gaps in how we regulate AI tools used for therapy and emotional support.

News
digital face mental health illustration

The Complexities of Governing Mental Health AI

Caroline Yee, Caroline Meinhardt, Michelle Mello, Jane Paik Kim
HealthcarePrivacy, Safety, SecurityGenerative AIRegulation, Policy, GovernanceJul 24

Policymakers, academics, healthcare providers, AI developers, and patient advocates convened by Stanford HAI identify critical gaps in how we regulate AI tools used for therapy and emotional support.

HAI Student Affinity Groups Take On Society’s Emerging Questions
Madeleine Wright
Jun 26, 2026
News

Stanford students across disciplines are teaming up to tackle society’s pressing questions in the age of AI.

News

HAI Student Affinity Groups Take On Society’s Emerging Questions

Madeleine Wright
Arts, HumanitiesGenerative AIEthics, Equity, InclusionPrivacy, Safety, SecurityJun 26

Stanford students across disciplines are teaming up to tackle society’s pressing questions in the age of AI.

Today's AI Talks Like “Nobody.” New Research Gives It Real Personality.
Jun 08, 2026
News
3D illustration of mirrored human profiles in blue and yellow layers

PsychAdapter lets researchers dial in on personality traits, age, and mental health characteristics to generate text that sounds like real individuals, opening the door to training simulations and personalized content.

News
3D illustration of mirrored human profiles in blue and yellow layers

Today's AI Talks Like “Nobody.” New Research Gives It Real Personality.

HealthcareGenerative AISciences (Social, Health, Biological, Physical)Jun 08

PsychAdapter lets researchers dial in on personality traits, age, and mental health characteristics to generate text that sounds like real individuals, opening the door to training simulations and personalized content.