
OpenAI releases MentalHealthBench, an open benchmark for AI responses in mental health conversations
OpenAI has released MentalHealthBench, an open benchmark built with more than 80 licensed mental health experts from 22 countries to evaluate how AI models respond in realistic mental health conversations.
OpenAI introduced MentalHealthBench on September 23, an open benchmark that measures how well AI systems respond in realistic mental health conversations. The company says the benchmark was co-created with more than 80 licensed psychologists and psychiatrists from 22 countries.
More than a billion people use ChatGPT every week, and they increasingly bring it personal topics: relationships, everyday stress and moments of crisis. OpenAI notes that most existing evaluations focused on emergency scenarios, leaving open the question of how models behave across the full range of everyday, less acute conversations.
What the benchmark covers
MentalHealthBench contains 1,215 synthetic conversations built with privacy-preserving techniques that mirror real usage patterns. Of those, 53.5% are non-acute, 18.2% are high-acuity and 28.3% involve emergencies. Adults make up 68.1% of the personas, teens aged 13–17 account for 21.2%, caregivers 4.9% and clinicians 5.8%.
Each conversation was reviewed by at least three experts, who wrote rubric criteria with weights from −10 to +10: positive points reward helpful behavior, negative ones penalize harmful responses. Only criteria agreed by at least two experts and not contradicted by a third were kept. An automated grader, GPT-5.6 Sol, scores model responses against them. The overall result breaks down into ten behavioral dimensions, including context-seeking, empathy, clinical accuracy and urgency calibration.
OpenAI evaluated several models on the benchmark. Reported results put GPT-6 Astra at 57.3%, GPT-6 Sol at 53.9% and Claude Opus 5.5 at 52.4%.
Experts versus users
OpenAI ran a separate study with 44 adults from 16 countries who had used AI for mental health or emotional support. Participants rated responses only to non-acute conversations, to avoid exposing them to distressing material. Users valued practical next steps and tone, while experts placed more emphasis on gathering relevant context and interpreting ambiguous situations. The company says user input complements expert judgment but does not replace it; the benchmark's final criteria remain based on expert consensus.
Why it matters
MentalHealthBench is released openly so researchers can examine the methods, run their own evaluations and build on the work, and so other developers can measure their models against the same rubric. OpenAI stresses that ChatGPT is not a substitute for therapy or professional care.
The company also says it has strengthened ChatGPT's responses in sensitive conversations, expanded access to crisis resources and added Trusted Contact. No benchmark captures everything that matters in a personal conversation, OpenAI acknowledges, but it hopes MentalHealthBench sets a higher standard.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.