Announcing Transluce's Mental Health Evaluation
Our evaluation of how 77 AI system variants respond to simulated users experiencing suicidal ideation, psychosis, and mania, covering 14 mental health-relevant behaviors.
Transluce carried out the most expansive independent evaluation to date of how leading AI models respond to users experiencing mental health crises, including suicidal ideation, psychosis, and mania. We outline our findings below, but in summary we found that newer models have significantly improved their responses to users in obvious crisis, while still sometimes struggling with less clear-cut behaviors like creative writing and roleplay about suicide. Through a collaboration with OpenAI and Anthropic, we received unique, anonymized insights into how real users chat with ChatGPT and Claude…
saved by
related reading
- GPT-5.6 Preview System Card - OpenAI Deployment Safety Hubdeploymentsafety.openai.com
- Today, we are releasing a research preview of our user model, along with a set of evaluations designed to measure how faithfully user models capture human behavior.persimmon.humansand.ai
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- User awareness in frontier modelstransluce.org
- GPT-6 Astra System Carddeploymentsafety.openai.com
- Translucetransluce.org
- AI Induced Psychosis: A shallow investigation — LessWronglesswrong.com
- gpt-4.pdfcdn.openai.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com
- bro can you imainge they literally dropped a new model while I'm writ… · tim-hua-01/ai-psychosis@dcfe240 · GitHubgithub.com
- Claude 4 System Cardwww-cdn.anthropic.com