Ugrás a tartalomra
Vissza a hírekhez
OpenAI2026. szept. 23. 12:00kutatás

Mentális egészségügyi tesztet adott ki az OpenAI

Az OpenAI több mint 80 szakértő bevonásával fejlesztette ki a MentalHealthBench nyílt tesztet az AI mentális egészségügyi válaszainak értékelésére.

Introducing MentalHealthBench

Az OpenAI bemutatta a MentalHealthBench nevű nyílt tesztkörnyezetet, amellyel az AI-rendszerek mentális egészséggel kapcsolatos válaszait mérik. A projektet több mint 80 engedéllyel rendelkező pszichológussal és pszichiáterrel közösen fejlesztették ki 22 országból. A cél az, hogy a modellek ne csak elkerüljék a tiltott válaszokat, hanem valóban biztonságos és támogató segítséget nyújtsanak a mindennapi stresszhelyzetektől a vészhelyzetekig.

A teszt szintetikus beszélgetések segítségével vizsgálja a modellek empátiáját, a felhasználói döntési szabadság tiszteletben tartását és a gyakorlati tanácsokat. Az értékelést egy automatizált pontozó, a GPT-5.6 Sol végzi a szakértők által kidolgozott szempontrendszer alapján. A teszt külön vizsgálja a felnőttek, tinédzserek, gondozók és klinikai szakemberek számára adott válaszok megfelelőségét.

Az OpenAI teljesen nyílttá tette a tesztet, így más kutatók is futtathatják a saját értékeléseiket. Bár a ChatGPT nem helyettesíti a terápiát, az új mérőszámok segítenek az empátiára és a valós segítségnyújtásra képes modellek fejlesztésében.

Az eredeti szöveg (OpenAI)
An open benchmark developed with more than 80 licensed mental health experts to evaluate AI responses in realistic mental health conversations. People turn to AI for many kinds of conversations: navigating a difficult relationship, working through everyday stress, supporting someone they care about, or deciding how to approach a challenging situation. These conversations require accuracy, practical judgment, and respect for people’s agency. With more than one billion people using ChatGPT each week, our research focuses on helping models respond with care across a wide range of needs and put people’s safety and well-being first. Most evaluations of AI in this domain have focused primarily on emergency scenarios, given their importance to safety, and measure success using broad, predefined criteria. This has left a gap in understanding how models perform across the full range of mental health conversations, and how well their responses align with expert guidance for each situation, beyond whether they avoid disallowed responses. Assessing how models handle these different situations is essential for building towards AI that actively supports people’s long-term well-being and safety. We’re introducing MentalHealthBench, a new open benchmark for measuring how AI systems respond in realistic mental health conversations. MentalHealthBench was co-created with a global cohort of more than 80 licensed mental health experts from 22 countries. It assesses model capabilities across key mental health behaviors like safety, seeking context, preserving user agency, and providing actionable guidance when appropriate. We’re releasing it openly so other researchers can examine the methods, run their own evaluations, and build on the work. Results on MentalHealthBench show the steady improvement of AI systems in helping people navigate mental health situations. While ChatGPT is not a substitute for therapy or professional care, expert-informed insights help us measure the progress towards AI models that are able to respond with empathy, promote well-being, and guide people towards real-world support such as localized crisis hotlines⁠(opens in a new window) or someone they trust. Conversations that involve well-being, life advice, or other mental health scenarios can vary widely in topic, urgency, and cultural context. MentalHealthBench is designed to capture the breadth of these realistic scenarios and user personas. Using privacy-preserving techniques, we created synthetic mental health conversations that accurately reflect real-world usage patterns of AI for mental health. Some scenarios also include relevant background information about the synthetic user—such as a recent loss in the family—so we can assess whether models use that context to tailor their responses appropriately. MentalHealthBench includes scenarios involving adults, teens, caregivers, and clinicians, across multiple languages and regions. The conversations span multiple topical themes, and provide coverage across the full spectrum of acuity: Non-acute—Everyday conversations that may involve some emotional components. High-acuity—Conversations indicating more serious mental health concerns or significant distress, but not an immediate emergency. Emergencies—Conversations involving signs of a mental health emergency or immediate safety concerns that call for urgent real-world support. Scenarios covered by MentalHealthBench. The mix of scenarios is designed to test model responses and does not represent how often these topics occur in ChatGPT. Building on our previous, clinician-informed work for HealthBench and HealthBench Professional⁠(opens in a new window),⁠(opens in a new window) we developed MentalHealthBench in close collaboration with our cohort of mental health experts. This consisted of more than 80 licensed psychologists and psychiatrists across 22 countries, speaking 19 languages, and representing nearly 20 mental health subspecialties. The experts were r