flâneur — a map of the web's best reading

Irretrievability; or, Murphy's Curse of Oneshotness upon ASI — LessWrong

lesswrong.com · 8,798 words · saved by 1 readers

Comment by Sam Marks - Lest the exegesis of my old comment continue, I'm happy to clarify my object-level view. I think that: * At each AI capability level, there is some probability of an irrecoverable catastrophe (e.g. AI killing or disempowering humanity). * * You could rephrase this as "There will be critical tries." * This probability is importantly sensitive to preparation that we do in advance using less capable AIs. * * This preparation includes things like alignment/control research on weaker systems, hardening the world, and work on extracting as much useful labor as possible (e.g. alignment research) out of weaker AI systems. * By "importantly" sensitive, I mean that if you try to forecast catastrophe risk without modeling the effect of preparatory work with weaker AIs, then your forecast will be substantially worse. * In particular, this means that I expect it is feasible in practice for humanity to do preparatory work with weaker AIs that substantially moves the overall probability of catastrophe. * Factors that influence the efficacy of this prior preparation include: how much time we have with the less capable AIs, whether we prioritize and execute well on the preparatory work, and how similar the less capable AIs are to the more capable ones (along certain relevant axes). * The above dynamic will only recur for a finite number of rounds before either a catastrophe occurs or we develop and hand off to AIs which will properly handle the situation from then on. In my view, Eliezer's writing on AI risk does a poor job of modeling the effect of preparatory work with weaker AI systems. While he often points out ways that some preparatory work with weaker AIs can fail to be useful, he doesn't often engage with IMO the most plausible reasons that some people expect it to be useful. I'd be shocked to learn that we agree on how useful preparatory work with weaker AIs will be. So it seems to me that there's an important part of his view that he hasn't explained well, an

x Irretrievability; or, Murphy's Curse of Oneshotness upon ASI — LessWrong World Modeling Curated 2026 Top Fifty: 22 % 367 Irretrievability; or, Murphy's Curse of Oneshotness upon ASI by Eliezer Yudkowsky 4th May 2026 26 min read 132 367 Example 1: The Viking 1 lander In the 1970s, NASA sent a pair of probes to Mars, the Viking 1 and Viking 2 missions. Total cost of $1B (1970), equivalent to about $7B (2025). The Viking 1 probe operated on Mars's surface for six years, before its battery began to seriously degrade. One might have thought a battery problem like that would spell the irrevocable

Explore this link on the map →

saved by

related reading