Skip to content
Let's talk Website revenue leak audit
News

OpenAI releases an open test for mental health chats

1 min read News · Models

MentalHealthBench scores how models answer in mental health conversations, using criteria written by 80+ psychologists and psychiatrists in 19 languages.

What is in it

On September 24 OpenAI released MentalHealthBench, an open evaluation set. It holds 1,215 synthetic but realistic mental health conversations and 5,262 grading criteria tied to them. More than 80 licensed psychologists and psychiatrists from 22 countries wrote the criteria in 19 languages. Of the conversations, 53.5% are non-acute, 18.2% high-acuity and 28.3% emergencies.

How it scores

Each criterion checks a single aspect of a response and carries a weight from minus 10 to plus 10. Helpful behaviour earns points and harmful behaviour loses them. The main behaviours measured are safety, asking for context, preserving the user's agency and giving actionable guidance when appropriate.

Results

In the first table GPT-6 Astra leads at 57.3%, followed by GPT-6 Sol (53.9%), Claude Opus 5.5 (52.4%) and GPT-6 Luna (50.2%). Older models scored lower: GPT-4o 32.1% and Gemini 2.5 Pro 29.5%. With even the best score under 60%, models still have a long way to go here. The set is open so other researchers can check the method and run their own evaluations.