✳flâneur — a map of the web's best reading
Why AI alignment could be hard with modern deep learning
cold-takes.com · 4,277 words · saved by 3 readers
Why would we program AI that wants to harm us? Because we might not know how to do otherwise.
Click lower right to download or find on Apple Podcasts, Spotify, Stitcher, etc. --> This is a guest post by my colleague Ajeya Cotra . Holden previously mentioned the idea that advanced AI systems (e.g. PASTA ) may develop dangerous goals that cause them to deceive or disempower humans. This might sound like a pretty out-there concern . Why would we program AI that wants to harm us? But I think it could actually be a difficult problem to avoid, especially if advanced AI is developed using deep learning (often used to develop state-of-the-art AI today). In deep learning, we don’t program a com
Explore this link on the map →saved by
related reading
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Thomas Larsen's Shortform — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- How likely is deceptive alignment? — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com