Chris Olah on X: "One of the ideas I find most useful from @AnthropicAI's Core Views on AI Safety post (https://t.co/Q9i2ujIbjm) is thinking in terms of a distribution over safety difficulty. Here's a cartoon picture I like for thinking about it: https://t.co/QYZBCTwHoo" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Post See new posts Conversation Chris Olah @ch402 One of the ideas I find most useful from @AnthropicAI 's Core Views on AI Safety post ( https:// anthropic.com/index/core-vie ws-on-ai-safety …) is thinking in terms of a distribution over safety difficulty. Here's a cartoon picture I like for thinking about it: ALT 5:31 PM · Jun 7, 2023 · 169.4K Views 19 121 662 277 Relevant View quotes Post your reply Reply Chris Olah @ch402 · Jun 7, 2023 A lot of AI safety discourse focuses on very specific models of AI and AI safety. These are interesting, but I don't know how I could be confident in any one. I prefer to accept that we're just very uncertain. One important axis of that uncertainty is roughly "difficulty". 1 1 53 6K Chris Olah @ch402 · Jun 7, 2023 In this lens, one can see a lot of safety research as "eating marginal probability" of things going well, progressively addressing harder and harder safety scenar
Chris Olah @ch402 One of the ideas I find most useful from @ AnthropicAI 's Core Views on AI Safety post ( anthropic.com/index/core-vie… ) is thinking in terms of a distribution over safety difficulty. Here's a cartoon picture I like for thinking about it: 4:31 PM · Jun 7, 2023 169.6K Views 19 0 1 9 95 0 9 5 657 0 6 5 7 275 0 2 7 5
Explore this link on the map →related reading
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Where I agree and disagree with Eliezer — LessWronglesswrong.com
- A Field Guide to AI Safety—Asteriskasteriskmag.com
- Shtetl-Optimized >> Blog Archive >> My AI Safety Lecture for UT Effective Altruismscottaaronson.blog
- AI Safety for Fleshy Humans: a whirlwind touraisafety.dance
- Critical review of Christiano's disagreements with Yudkowsky — LessWronglesswrong.com
- Which side of the AI safety community are you in? — LessWronglesswrong.com
- AI safety - Wikipediaen.wikipedia.org
- AI #24: Week of the Podcast — LessWronglesswrong.com
- I'm Switching Into AI Safetyalexirpan.com