✳flâneur — a map of the web's best reading
Alignment Is Proven To Be Solvable - by SE Gyges
verysane.ai · 3,194 words · saved by 1 readers
That LLMs understand natural language as well as they do should dramatically change our understanding of the problem.
Alignment Is Proven To Be Solvable SE Gyges Feb 18, 2026 28 5 6 Share At least the systems that we build today often have that property. I mean, I’m hopeful that someday we’ll be able to build systems that have more of a sense of common sense. We talk about possible ways to address this problem, but yeah I would say it is like this Genie problem. Dario Amodei, Concrete Problems in AI Safety with Dario Amodei and Seth Baum , 2016 We might call this the King Midas problem: Midas, a legendary king in ancient Greek mythology, got exactly what he asked for—namely, that everything he touched should
Explore this link on the map →saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- the void — LessWronglesswrong.com
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Alignment will happen by default. What’s next? — LessWronglesswrong.com
- What Is The Alignment Problem? — LessWronglesswrong.com
- Alignment By Default — AI Alignment Forumalignmentforum.org