AI Safety Seems Hard to Measure
cold-takes.com · 5,816 words · saved by 2 readers
Four analogies for why "We don't see any misbehavior by this AI" isn't enough.
Click lower right to download or find on Apple Podcasts, Spotify, Stitcher, etc. In previous pieces, I argued that there's a real and large risk of AI systems' developing dangerous goals of their own and defeating all of humanity - at least in the absence of specific efforts to prevent this from happening. A young, growing field of AI safety research tries to reduce this risk, by finding ways to ensure that AI systems behave as intended (rather than forming ambitious aims of their own and deceiving and manipulating humans as needed to accomplish them). Maybe we'll succeed in reducing the risk,
saved by
related reading
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- AI #24: Week of the Podcast — LessWronglesswrong.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- What failure looks like — LessWronglesswrong.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- How we could stumble into AI catastrophecold-takes.com
- AI Safety for Fleshy Humans: a whirlwind touraisafety.dance