TREE NEWS update: OpenAI has released MentalHealthBench, an open-source benchmark built with the participation of more than 80 mental health clinicians, to evaluate frontier models across real-world mental health conversations. The company says the benchmark spans the full range of mental health dialogue, from everyday psychological support to more severe crisis scenarios, unlike most existing benchmarks that focus on emergencies. OpenAI published it for other researchers to inspect the methods, run their own evaluations and build on the work.
OpenAI Open-Sources MentalHealthBench Built With 80+ Clinicians
The notable move here is governance, not capability: a frontier lab is handing external researchers the evaluation methodology itself, not just a scoreboard, which lets third parties audit how mental health performance is defined and where models fail. That matters most for clinicians and researchers, who currently lack shared standards for judging these systems on everyday support rather than only crisis response. The open question is whether other labs adopt or contest the benchmark, and whether clinician-designed evaluations become the default reference point for this category.
Generated by AI for reference only.
Share on WeChat
Open WeChat → Scan → then tap "…" to send to a chat or Moments.
Tap "…" in the top-right corner to send to a chat or share to Moments.