The Complexities of Governing Mental Health AI

Policymakers, academics, healthcare providers, AI developers, and patient advocates convened by Stanford HAI identify critical gaps in how we regulate AI tools used for therapy and emotional support.
Millions of Americans are affected by mental illness every year, yet the cost of therapy remains out of reach for many, and a shortage of licensed clinicians means that even those with insurance often wait months for an appointment. Into this void has stepped a rapidly expanding market of AI-powered tools, including chatbots that provide therapeutic counseling and apps that offer cognitive behavioral therapy exercises on demand. Children and adults also seek out general-purpose chatbots like ChatGPT and “companion” bots such as those offered by Character.ai and Replika in times of loneliness or emotional distress.
There are promising potential upsides to the use of AI in mental health care: greater access, lower cost, tools that could extend the reach of an overstretched system, and a form of social and emotional support for people experiencing loneliness. But these promises also entail risk. Absent clear regulation and standardized third-party testing, these tools risk delivering substandard care and putting users in danger. News headlines abound about minors developing unhealthy emotional attachments to chatbots, users in crisis receiving harmful or inadequate responses, and research showing that general-purpose AI chatbots commonly miss warning signs.
AI’s role in mental health care is growing fast, and legislators are struggling to keep pace. To date, most legislative activities have happened in states, which have introduced more than 140 bills related to AI use in mental health contexts. Federal legislation is pending, but so far has been narrowly focused on protection of minors. This fragmented policy landscape is further hindered by a perpetually lagging evidence base: Many purpose-built AI mental health tools lack validated outcomes and representative samples, and are rarely evaluated with rigorous study designs. Meanwhile, the models powering general-purpose chatbots update so rapidly that safety research findings quickly become outdated.
Recognizing these governance challenges, the Stanford Institute for Human-Centered AI (HAI) convened a select group of leading researchers, clinicians, policymakers, behavioral health experts, ethicists, AI developers, and patient advocates for a policy workshop on mental health and AI in June 2026. The meeting, which followed the inaugural AI for Mental Health Symposium held earlier that day, was hosted by HAI’s Healthcare AI Policy Steering Committee in collaboration with the university’s AI for Mental Health (AI4MH) Initiative.
Under the Chatham House Rule, participants had candid discussions about emerging efforts to regulate the use of AI in mental health care; evidentiary gaps that must be addressed to enable sound policymaking; and the technical feasibility of potential policy levers and guardrails. Below, we summarize three key policy challenges this group identified for further research.
1. The Field Needs Clearer Definitions
Effective regulation will require policymakers, mental health practitioners, and AI developers to agree on the boundaries of “mental health AI” – a broad umbrella term that can refer to many different tools and applications (see table). Right now, there is no broad consensus on what specifically counts as an AI mental health tool, where the line falls between a clinical function that needs to be regulated and a wellness feature, or which policy levers to apply to which products. Should a chatbot that was clinically developed specifically for mental health contexts be treated the same as a general-purpose chatbot not originally designed for those purposes but that a user turns to in a mental health crisis?
While all mental health AI tools should meet some baseline safety expectations, different types of AI-driven mental health tools may create different expectations and risks, and thus require different regulation. Yet many legislative approaches are not sensitive to the differentiated implications of various mental health AI products. For example, outright bans on all AI used by human providers to deliver psychotherapy services don’t address the reality that people may still turn to general-purpose chatbots for therapy that is not yet vetted or supervised by clinicians. In other words, a law meant to keep AI out of therapeutic relationships may instead push users toward less tailored and unregulated general tools. Until there is definitional clarity, regulation will remain fragmented, evaluation standards will be inconsistent, and companies will continue to operate in ambiguity.
An Illustrative Overview of Mental Health AI Categories
Type | Description |
General-purpose LLMs | Examples: ChatGPT, Claude, Gemini AI tools that are not specifically designed for mental health purposes, but are commonly used for mental health support |
Companion chatbots | Examples: ChatGPT’s CounselorGPT, Character.ai’s Trauma Therapist, and Meta AI Studio’s My Therapist General-purpose LLMs that assume a particular persona or role, such as that of a therapist, friend, or trusted partner |
Wellness apps that use AI | Examples: Wysa Applications that use AI to provide mental health support and promote emotional well-being more broadly but stop short of making medical claims – and therefore fall outside the regulatory scope of medical devices |
Purpose-built LLMs for mental health support | Examples: TheraBot, Slingshot.AI’s Ash, TalkSpace’s Tee LLMs developed specifically for application in mental health care, usually developed through clinical testing but varying in quality |
AI tools used in non-treatment mental health care settings | Examples: Mentalyc, Commure, TherapyNotes AI tools used by human therapists to assist with administrative work (e.g., clinical notetaking), training (e.g., upskilling novice counselors), or case management (e.g., resource or benefits navigation). These uses can still affect mental health outcomes and risk, even if they do not alone constitute psychotherapy. |
2. Evaluation Methods Are Lagging Behind
Although methods for systematically evaluating the performance of other healthcare AI tools are rapidly developing, methods to assess how effectively and safely mental health AI tools function in the real world are lagging behind. The highest-stakes interactions are rare and hard to simulate, model behavior is unpredictable in real-world situations, and single-session testing reveals little about long-term effects. Meanwhile, the real-world chat data needed to study safety and efficacy at scale sits largely within industry, with no meaningful structures for sharing with independent researchers or regulators.
This creates a policy bottleneck. Many reasonable safety requirements can’t be justified or enforced without large-scale interaction data. Consider sycophancy: the tendency of AI tools to validate and affirm users, which may be particularly harmful for people with OCD, where validation-seeking and prolonged engagement can reinforce compulsive patterns. Developing and enforcing suitable policy mechanisms to address the issue requires knowing how often it happens, to whom, and with what effect. We currently can’t answer those questions.
Even where data exists, the evaluation landscape is fragmented. While researchers have proposed many different approaches to benchmarking, there is a lack of consensus on what is most important to measure and how, and whether benchmarking is even the right tool for systems whose behavior changes week to week. Of the many already existing metrics, most reflect technical priorities set by developers (e.g., percentage of messages indicating possible signs of mental health emergencies) rather than the goals and techniques of mental health care (e.g., patient-tailored treatments), which are themselves contested. Most evaluation frameworks currently focus on single, point-in-time analysis and thus are poorly suited to capturing how chatbot use affects users long term. This is a significant gap: Early evidence suggests prolonged use of some types of chatbots may worsen well-being, making longitudinal assessment essential. Policymakers, researchers, and industry must work further together to standardize evaluations of mental health AI and move them toward multimodal assessments that span multiple back-and-forth chatbot exchanges.
3. Target Low-Hanging Fruit While Tackling Deeper Issues
Workshop participants agreed that more comprehensive mental health AI policy is urgently needed. The good news is that there are some low-hanging fruit that policymakers can and should act on quickly. Transparency requirements, crisis response mechanisms, data protection measures, and parental controls for minors have broad agreement and urgency, which is why they’re already among the most commonly passed provisions in state mental health AI law. When a state enacts thoughtful regulations on such issues, it can set a precedent for other states and lead to some meaningful safety improvements.
However, long-term, meaningful change will require resolving deeper policy challenges and resolving tensions across today’s patchwork of state laws. Disparate state laws concerning therapy tools can be onerous for psychotherapists who are licensed in multiple states with different regulations. Perhaps the most underexamined and unresolved issue in this conversation is one of the most fundamental: Business models built around maximizing user engagement are structurally at odds with the goal of fostering a healthy relationship with chatbots. We have watched this play out before – courts have already tied social media’s “keep them on the platform” logic to addiction and harm to minors, and research links heavy use to worse mental health outcomes and suicidal behavior. A chatbot optimized to simulate an intimate human relationship, too, can foster overuse and overreliance. Without mechanisms that reward responsible behavior, there is little reason to expect the industry to self-regulate differently.
Finally, the policy conversation is currently too narrow. It is dominated by higher-income, commercially insured, and professionally licensed perspectives, with less representation of people with severe mental illness, young people, and individuals in the social welfare and criminal justice systems. Leaving them out risks compounding existing inequities at scale, especially at a time when policymakers are urgently seeking fast remedies.





