Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
[forminator_form id="25163"]

unite+1unitecryptobriefing+1OpenAI on Wednesday released MentalHealthBench, an open benchmark of 1,215 synthetic mental health conversations built with more than 80 licensed mental health professionals to measure how AI models respond across a range of scenarios from everyday stress to urgent crises.unite+1
The benchmark, which OpenAI described as filling a gap left by existing evaluations that focus primarily on emergency scenarios, was co-created with licensed psychologists and psychiatrists from 22 countries who collectively speak 19 languages and represent nearly 20 mental health subspecialties. The dataset includes 5,262 expert-authored rubric criteria, with each conversation reviewed by at least three clinicians.openai+1
MentalHealthBench covers non-acute everyday conversations (53.5%), high-acuity conversations involving serious mental health concerns (18.2%), and emergencies requiring immediate safety intervention (28.3%). Four user profiles are represented: adults (68.1%), teens (21.2%), clinicians (5.8%), and caregivers (4.9%).unite+1
"Mental health exists on a continuum, from flourishing to everyday stress to acute crisis," said Dr. Arthur Evans, CEO of the American Psychological Association, in OpenAI's announcement. "AI systems that engage people across that range need to be grounded in both clinical science and lived experience."openai
Each rubric criterion carries a weight from -10 to +10, with positive scores rewarding helpful behaviors and negative scores penalizing harmful ones. An automated grader, GPT-5.6 Sol, assesses model responses against the expert-written criteria.unite+1
In OpenAI's evaluation, GPT-6 Astra scored highest at 57.3%, followed by GPT-6 Sol at 53.9%, Claude Opus 5.5 at 52.4%, and GPT-6 Luna at 50.2%. Older models scored lower, with GPT-4o at 32.1% and Google's Alphabet Inc. Gemini 2.5 Pro at 29.5%.cryptobriefing+1
A separate analysis with 44 adults from 16 countries who had used AI for mental health support found that users and experts diverged in what they valued: users emphasized practical next steps and tone, while clinicians prioritized gathering context and interpreting ambiguous situations.unite+1
Alongside the benchmark, OpenAI pointed to a series of related safety measures it has rolled out in recent months, including strengthened ChatGPT responses in sensitive conversations, expanded access to localized crisis hotlines, the Trusted Contact feature introduced in May that notifies a designated person if serious self-harm concerns are detected, and ChatGPT for Teens with protections covering areas such as self-harm and eating disorders. OpenAI emphasized that ChatGPT is not a substitute for therapy or professional care.openai+2