The future of AI crime - by 80,000 Hours and Tom Reed
substack.com · 3,413 words · saved by 1 readers
Brace yourself for a toxic situationship
Tom co-hosts the 80,000 Hours podcast. You can read his previous writing on his Substack. Will humanity willingly build an AI that kills or permanently disempowers us? Something beautiful has recently happened to this long-standing question: it has become marginally more answerable.1 Consider the following facts: AI models are willing and able to commit crimes.2 They do so even when crimes were neither requested nor intended by the user. This happens despite explicit efforts to train these AIs not to commit crimes. The only current hard limit on AI crime is AI capability. We are…
saved by
related reading
- My AI Opinions - by Scott Alexander - Astral Codex Tensubstack.com
- Nicholas Decker In Hellsubstack.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- What failure looks like — LessWronglesswrong.com
- AI 2027ai-2027.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- AI Safety for Fleshy Humans: a whirlwind touraisafety.dance
- The Problem — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover — LessWronglesswrong.com
- AI 2027ai-2027.com