flâneur — a map of the web's best reading

Shah and Yudkowsky on alignment failures — LessWrong

lesswrong.com · 27,676 words · saved by 1 readers

I am a time-traveler who came back from the world where it (super duper predictably) turned out that a lot of early bright hopes didn't pan out and various things went WRONG and alignment was HARD and it was NOT SOLVED IN ONE SUMMER BY TEN SMART RESEARCHERS I think these kinds of comments update readers' beliefs in a bad, invalid way. The bad event (AGI ruin) is argued for by... a request for me to condition on testimony of a survivor of that bad event. Yes, I know the whole thing is tongue-in-cheek. I know that EY is not literally claiming to be a time-traveller. But in TurnTrout-culture, "I experienced X" is something to be said when X has actually been experienced. "The fact that X" is to be said when X is actually supported by a heap of accepted evidence.[1] "Have you met dath ilani?" is to be said when such entities actually exist and are not outputs of the model of intelligence which is being argued for. (Yes, that last one was flagged as a "bad argument", but still.) This paragr

x Shah and Yudkowsky on alignment failures — LessWrong 2021 MIRI Conversations AI Frontpage 93 Shah and Yudkowsky on alignment failures by Rohin Shah , Eliezer Yudkowsky 28th Feb 2022 AI Alignment Forum 110 min read 47 93 Ω 41 This is the final discussion log in the Late 2021 MIRI Conversations sequence, featuring Rohin Shah and Eliezer Yudkowsky, with additional comments from Rob Bensinger, Nate Soares, Richard Ngo, and Jaan Tallinn. The discussion begins with summaries and comments on Richard and Eliezer's debate. Rohin's summary has since been revised and published in the Alignment Newslett

Explore this link on the map →

related reading