My personal cruxes for working on AI safety - EA Forum
The following is a heavily edited transcript of a talk I gave for the Stanford Effective Altruism club on 19 Jan 2020. I had rev.com transcribe it, and then Linchuan Zhang, Rob Bensinger and I edited it for style and clarity, and also to occasionally have me say smarter things than I actually said. Linch and I both added a few notes throughout. Thanks also to Bill Zito, Ben Weinstein-Raun, and Howie Lempel for comments. I feel slightly weird about posting something so long, but this is the natural place to put it. Over the last year my beliefs about AI risk have shifted moderately; I expect that in a year I'll think that many of the things I said here were dumb. Also, very few of the ideas here are original to me. -- After all those caveats, here's the talk: INTRODUCTION It's great to be here. I used to hang out at Stanford a lot, fun fact. I moved to America six years ago, and then in 2015, I came to Stanford EA every Sunday, and there was, obviously, a totally different crop of people there. It was really fun. I think we were a lot less successful than the current Stanford EA iteration at attracting new people. We just liked having weird conversations about weird stuff every week. It was really fun, but it's really great to come back and see a Stanford EA which is shaped differently. Today I'm going to be talking about the argument for working on AI safety that compels me to work on AI safety, rather than the argument that should compel you or anyone else. I'm going to try to spell out how the arguments are actually shaped in my head. Logistically, we're going to try to talk for about an hour with a bunch of back and forth and you guys arguing with me as we go. And at the end, I'm going to do miscellaneous Q and A for questions you might have. And I'll probably make everyone stand up and sit down again because it's unreasonable to sit in the same place for 90 minutes. META LEVEL THOUGHTS I want to first very briefly talk about some concepts I have that a
adamShimi 6y 20 0 0 Thanks a lot for this great post! I think the part I like the most, even more than the awesome deconstruction of arguments and their underlying hypotheses, is the sheer number of times you said "I don't know" or "I'm not sure" or "this might be false". I feel it places you at the same level than your audience (including me), in the sense that you have more experience and technical competence than the rest of us, but you still don't know THE TRUTH, or sometimes even good approximations to it. And the standard way to present clearly ideas and research is to structure them so
Explore this link on the map →related reading
- Where I agree and disagree with Eliezer — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Why I don't persuade people to do AI safetyjason.ml
- AI safety undervalues founders — LessWronglesswrong.com
- AGI safety career advice — EA Forumforum.effectivealtruism.org
- I'm Switching Into AI Safetyalexirpan.com
- AGI safety career advice — LessWronglesswrong.com
- "Taking AI Risk Seriously" (thoughts by Critch) — LessWronglesswrong.com
- Critical review of Christiano's disagreements with Yudkowsky — LessWronglesswrong.com
- Benefits & Risks of Artificial Intelligence - Future of Life Institutefutureoflife.org
- AI safety technical research | Career review | 80,000 Hours80000hours.org
- A Field Guide to AI Safety—Asteriskasteriskmag.com