Many arguments for AI x-risk are wrong — AI Alignment Forum
The following is a lightly edited version of a memo I wrote for a retreat. It was inspired by a draft of Counting arguments provide no evidence for AI doom. I think that my post covers important points not made by the published version of that post. I'm also thankful for the dozens of interesting conversations and comments at the retreat. I think that the AI alignment field is partially founded on fundamentally confused ideas. I’m worried about this because, right now, a range of lobbyists and concerned activists and researchers are in Washington making policy asks. Some of these policy proposals seem to be based on erroneous or unsound arguments.[1] The most important takeaway from this essay is that the (prominent) counting arguments for “deceptively aligned” or “scheming” AI provide ~0 evidence that pretraining + RLHF will eventually become intrinsically unsafe. That is, that even if we don't train AIs to achieve goals, they will be "deceptively aligned" anyways. This has important
x Many arguments for AI x-risk are wrong — AI Alignment Forum Deceptive Alignment AI Risk Skepticism AI Governance Counting arguments Language Models (LLMs) AI Community Frontpage 48 Many arguments for AI x-risk are wrong by TurnTrout 5th Mar 2024 15 min read 96 48 The following is a lightly edited version of a memo I wrote for a retreat. It was inspired by a draft of Counting arguments provide no evidence for AI doom . I think that my post covers important points not made by the published version of that post. I'm also thankful for the dozens of interesting conversations and comments at the r
saved by
related reading
- Many arguments for AI x-risk are wrong — LessWronglesswrong.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Unfalsifiable stories of doom | Mechanize, Inc.mechanize.work
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AI in 2025: gestalt — LessWronglesswrong.com
- How will we update about scheming?blog.redwoodresearch.org
- Existential Risk from AI: An Exposition for Mathematiciansalkjash.github.io
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- The Best of LessWrong — LessWronglesswrong.com
- Why AI alignment could be hard with modern deep learningcold-takes.com