Many arguments for AI x-risk are wrong — AI Alignment Forum
The following is a lightly edited version of a memo I wrote for a retreat. It was inspired by a draft of Counting arguments provide no evidence for AI doom. I think that my post covers important points not made by the published version of that post. I'm also thankful for the dozens of interesting conversations and comments at the retreat. I think that the AI alignment field is partially founded on fundamentally confused ideas. I’m worried about this because, right now, a range of lobbyists and concerned activists and researchers are in Washington making policy asks. Some of these policy proposals seem to be based on erroneous or unsound arguments.[1] The most important takeaway from this essay is that the (prominent) counting arguments for “deceptively aligned” or “scheming” AI provide ~0 evidence that pretraining + RLHF will eventually become intrinsically unsafe. That is, that even if we don't train AIs to achieve goals, they will be "deceptively aligned" anyways. This has important
x Many arguments for AI x-risk are wrong — AI Alignment Forum Deceptive Alignment AI Risk Skepticism AI Governance Counting arguments Language Models (LLMs) AI Community Frontpage 48 Many arguments for AI x-risk are wrong by TurnTrout 5th Mar 2024 15 min read 96 48 The following is a lightly edited version of a memo I wrote for a retreat. It was inspired by a draft of Counting arguments provide no evidence for AI doom . I think that my post covers important points not made by the published version of that post. I'm also thankful for the dozens of interesting conversations and comments at the r
Explore this link on the map →related reading
- Many arguments for AI x-risk are wrong — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Unfalsifiable stories of doom | Mechanize, Inc.mechanize.work
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- How will we update about scheming? — LessWronglesswrong.com
- The Best of LessWrong — LessWronglesswrong.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org