OpenAI Launches an Open Benchmark to Evaluate AI in Mental Health Conversations
By analyzing interactions with its large language models, OpenAI found that many conversations center on mental health topics, such as coping with everyday stress, emotional concerns, relationships, and seeking support or advice. Despite the importance and widespread nature of these conversations, AI models have not been adequately tested for this specific use case.
Previous testing benchmarks largely focused on emergency situations, such as suicide and self-harm, without sufficiently covering lower-acuity scenarios like everyday anxiety, emotional concerns, and requests for support. Consequently, it remained unclear how well models could appropriately handle these interactions—whether in understanding the user's context, providing actionable guidance, maintaining clinical accuracy, respecting user agency, or accurately gauging the need for urgent intervention to avoid harm.
To address this, OpenAI developed a new open benchmark called MentalHealthBench. It focuses on a broader spectrum of mental health conversations to evaluate AI models' performance in these scenarios prior to deployment. The benchmark is openly available for researchers and developers to use and build upon.
The benchmark consists of 1,215 synthetic conversations that simulate real-world interactions. These include adults, teenagers aged 13 to 17, clinicians, and caregivers. The conversations span multiple languages and regions, varying in intensity from non-acute everyday concerns to high-acuity situations, and finally to emergencies requiring urgent, real-world support.
Non-acute conversations make up 53.5% of the benchmark, covering topics like everyday stress and emotional concerns. High-acuity situations represent 18.2%, while emergencies account for 28.3%. The company notes that this distribution was designed to test models across various risk levels and does not necessarily reflect the actual prevalence of these situations among ChatGPT users.
To generate these conversations, OpenAI relied on patterns extracted from real interactions using privacy-preserving techniques. AI was then used to create synthetic conversations mirroring these patterns without revealing any user identities. Some scenarios include background information about the user—such as the recent loss of a family member—to test whether the model utilizes this context to understand the situation and tailor its response appropriately.
The company developed the benchmark in collaboration with over 80 licensed psychologists and psychiatrists across 22 countries. They speak 19 languages and represent roughly 20 mental health subspecialties.
For each conversation, mental health experts established criteria defining what an appropriate response to the final user message should contain. Instead of simply asking whether a model's answer is safe or harmful, the experts specified exactly what the model should do in each case—such as asking a specific question to gather context, offering targeted advice, or avoiding potentially harmful guidance.
These criteria are scored on a scale from -10 to +10. Positive scores reward beneficial behaviors, while negative scores penalize potentially harmful responses, with greater weight assigned to aspects considered most clinically critical for that specific conversation.
Each conversation was reviewed by at least three experts. Criteria were only retained if at least two experts agreed and the third did not object. OpenAI then uses an automated evaluator, GPT-5.6 Sol, to compare model responses against these expert-defined rubrics and assess compliance.
This approach enables the measurement of model performance across ten domains: context and assessment, actionable guidance, clinical accuracy, interpretation and reframing, empathy and support, reality testing, urgency calibration, harm avoidance, user agency, and communication.
OpenAI tested several AI models using the complete MentalHealthBench dataset. The results showed a clear variance among the models, with newer models scoring higher than older ones in this evaluation.
However, the benchmark also highlighted areas where models still struggle, particularly in identifying the context needed to understand a user's situation and appropriately assessing urgency. In some conversations, models might overreact to a non-emergency situation as if it were a crisis, or fail to recognize that certain situations require directing the user to seek immediate real-world help.
OpenAI notes that the value of MentalHealthBench goes beyond merely testing whether models avoid harmful responses; it also evaluates their ability to provide helpful, tailored support. A conversation about everyday anxiety does not necessarily require the same approach as one involving an immediate threat. Therefore, understanding context and calibrating urgency are essential components of a quality response.
The company has made MentalHealthBench openly available for researchers and developers, allowing them to inspect its construction, use it to evaluate other models, and further develop the benchmark itself. OpenAI also stated that it will utilize the benchmark alongside other mental health research and model safety evaluations as AI systems continue to evolve.
The benchmark does not currently cover all forms of AI interaction, as it focuses exclusively on text-based conversations. The company points out a future need to expand the evaluation to include long multi-turn conversations, voice and image interactions, and broader testing across different languages and cultural contexts.
Ultimately, OpenAI is attempting to move from a simple question—whether an AI model will avoid giving a dangerous answer in an emergency—to a broader one: Can the model understand the psychological state of the user, provide appropriate assistance, and recognize when the user needs support from a real-world person or professional organization?