Journalism begins where hype ends

,,

The danger of AI is not that it will become conscious and hate us, but that it will become competent and ignore us."

—Eliezer Yudkowsky

MentalHealthBench: OpenAI’s Test for AI in Everyday Distress and Crisis Talks

OpenAI released MentalHealthBench on Sept. 23 to measure how AI models handle realistic mental health talks, from daily stress to crises. The move comes as the company faces multiple U.S. lawsuits from families who say loved ones confided in ChatGPT before suicide or, in one case, a homicide.
OpenAI logo next to a human head silhouette with a tree inside representing thoughts and mental health, used with MentalHealthBench announcement.
September 25, 2026 11:34 AM IST | Written by Vaibhav Jha

As people increasingly turn to AI chatbots to seek mental health advice, OpenAI, the creators of ChatGPT, has introduced MentalHealthBench, an open benchmark to evaluate AI models on realistic conversations with users from their everyday feelings to crisis scenarios.

Of late, OpenAI has found itself facing multiple lawsuits in US, concerning wrongful deaths and in one case, a homicide, where the families have alleged that the victims/perpetrators had extended conversations with ChatGPT and the parent company failed to notify to authorities in crisis time.

In the light of these lawsuits and increasing use of AI chatbots to seek mental health advice, the MentalHealthBench introduced by OpenAI gives a platform to check whether AI models can replicate the advice/suggestions of a mental health professional.

According to OpenAI, the benchmark has been co-developed with over 80 licenses psychologists and psychiatrists across 22 countries, speaking 19 languages and representing 20 mental health subspecialities.

The benchmark contains 1,215 synthetic dialogues paired with 5,262 rubric items, with non-acute everyday exchanges making up 53.5 percent of the set, high-acuity distress 18.2 percent and emergencies 28.3 percent.

“The benchmark assesses model capabilities across key mental health behaviors like safety, seeking context, preserving user agency, and providing actionable guidance when appropriate. We’re releasing it openly so other researchers can examine the methods, run their own evaluations, and build on the work,” read an excerpt from OpenAI.

According to OpenAI, dialogues were generated with privacy-preserving methods to reflect actual ChatGPT patterns and models were scored against the expert rubrics by an automated grader.

In the benchmark, GPT-6 Astra led with 57.3 percent, followed by GPT-6 Sol at 53.9 percent and Claude Opus 5.5 at 52.4 percent. GPT-4o, tested from March 2025, scored 32.1 percent and Gemini 2.5 Pro 29.5 percent, claimed OpenAI’s report.

OpenAI said the results show steady gains across model generations while noting remaining gaps in gathering context and matching urgency to the situation. The company stated ChatGPT is not a substitute for therapy or professional care and is intended to steer users toward crisis lines or trusted people.

Also Read: AI as Therapist? Here’s Why That’s Not a Good Idea

Author

  • Vaibhav Jha, editor and co-founder at AI FrontPage

    Vaibhav Jha is an Editor and Co-founder of AI FrontPage. In his decade long career in journalism, Vaibhav has reported for publications including The Indian Express, Hindustan Times, and The New York Times, covering the intersection of technology, policy, and society. Outside work, he’s usually trying to persuade people to watch Anurag Kashyap films.

    LinkedIn