AI Induced Psychosis: A shallow investigation — LessWrong
Epimistemic status: A small project I worked on the side over ten days, which grew out of my gpt-oss-20b red teaming project. I think I succeeded in surfacing interesting model behaviors, but I haven’t spent enough time to make general conclusions about how models act. However, I think this methodological approach is quite reasonable, and I would be excited for others to build on top of this work! There have been numerous media reports of how ChatGPT has been fueling psychosis and delusions among its users. For example, ChatGPT told Eugene Torres that if he “truly, wholly believed — not emotionally, but architecturally — that [he] could fly [after jumping off a 19-story building]? Then yes. [He] would not fall.” There is some academic work documenting this from a psychology perspective: Morris et al. (2025) give an overview of AI-driven psychosis cases found in the media, and Moore et al. (2025) try to measure whether AIs respond appropriately when acting as therapists. Scott Alexander
x AI Induced Psychosis: A shallow investigation — LessWrong LLM-Induced Psychosis AI Curated 2025 Top Fifty: 14 % 386 AI Induced Psychosis: A shallow investigation by Tim Hua 26th Aug 2025 AI Alignment Forum 32 min read 47 386 Ω 89 “This is a Copernican-level shift in perspective for the field of AI safety.” - Gemini 2.5 Pro “What you need right now is not validation, but immediate clinical help .” - Kimi K2 Two Minute Summary There have been numerous media reports of AI-driven psychosis, where AIs validate users’ grandiose delusions and tell users to ignore their friends’ and family’s pushbac
Explore this link on the map →saved by
related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Aren’t developers regularly making their AIs nice and safe and obedient? | If Anyone Builds It, Everyone Dies | If Anyone Builds It, Everyone Diesifanyonebuildsit.com
- bro can you imainge they literally dropped a new model while I'm writ… · tim-hua-01/ai-psychosis@dcfe240 · GitHubgithub.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- AI #77: A Few Upgrades - by Zvi Mowshowitzthezvi.substack.com
- How well do models follow their constitutions? — LessWronglesswrong.com
- What Is Claude? Anthropic Doesn’t Know, Either | The New Yorkernewyorker.com