flâneur — a map of the web's best reading

Many arguments for AI x-risk are wrong — AI Alignment Forum

alignmentforum.org · 18,635 words · saved by 1 readers

The following is a lightly edited version of a memo I wrote for a retreat. It was inspired by a draft of Counting arguments provide no evidence for AI doom. I think that my post covers important points not made by the published version of that post. I'm also thankful for the dozens of interesting conversations and comments at the retreat. I think that the AI alignment field is partially founded on fundamentally confused ideas. I’m worried about this because, right now, a range of lobbyists and concerned activists and researchers are in Washington making policy asks. Some of these policy proposals seem to be based on erroneous or unsound arguments.[1] The most important takeaway from this essay is that the (prominent) counting arguments for “deceptively aligned” or “scheming” AI provide ~0 evidence that pretraining + RLHF will eventually become intrinsically unsafe. That is, that even if we don't train AIs to achieve goals, they will be "deceptively aligned" anyways. This has important

x Many arguments for AI x-risk are wrong — AI Alignment Forum Deceptive Alignment AI Risk Skepticism AI Governance Counting arguments Language Models (LLMs) AI Community Frontpage 48 Many arguments for AI x-risk are wrong by TurnTrout 5th Mar 2024 15 min read 96 48 The following is a lightly edited version of a memo I wrote for a retreat. It was inspired by a draft of Counting arguments provide no evidence for AI doom . I think that my post covers important points not made by the published version of that post. I'm also thankful for the dozens of interesting conversations and comments at the r

Explore this link on the map →

related reading