OpenAI releases an open test for mental health chats
MentalHealthBench scores how models answer in mental health conversations, using criteria written by 80+ psychologists and psychiatrists in 19 languages.
What is in it
On September 24 OpenAI released MentalHealthBench, an open evaluation set. It holds 1,215 synthetic but realistic mental health conversations and 5,262 grading criteria tied to them. More than 80 licensed psychologists and psychiatrists from 22 countries wrote the criteria in 19 languages. Of the conversations, 53.5% are non-acute, 18.2% high-acuity and 28.3% emergencies.
How it scores
Each criterion checks a single aspect of a response and carries a weight from minus 10 to plus 10. Helpful behaviour earns points and harmful behaviour loses them. The main behaviours measured are safety, asking for context, preserving the user's agency and giving actionable guidance when appropriate.
Results
In the first table GPT-6 Astra leads at 57.3%, followed by GPT-6 Sol (53.9%), Claude Opus 5.5 (52.4%) and GPT-6 Luna (50.2%). Older models scored lower: GPT-4o 32.1% and Gemini 2.5 Pro 29.5%. With even the best score under 60%, models still have a long way to go here. The set is open so other researchers can check the method and run their own evaluations.