Alignment Is Proven To Be Solvable - by SE Gyges
verysane.ai · 3,194 words · saved by 3 readers
That LLMs understand natural language as well as they do should dramatically change our understanding of the problem.
Alignment Is Proven To Be Solvable SE Gyges Feb 18, 2026 28 5 6 Share At least the systems that we build today often have that property. I mean, I’m hopeful that someday we’ll be able to build systems that have more of a sense of common sense. We talk about possible ways to address this problem, but yeah I would say it is like this Genie problem. Dario Amodei, Concrete Problems in AI Safety with Dario Amodei and Seth Baum , 2016 We might call this the King Midas problem: Midas, a legendary king in ancient Greek mythology, got exactly what he asked for—namely, that everything he touched should
saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- The Artificiality of Alignmentjoinreboot.org
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- the void — LessWronglesswrong.com
- Training language models to follow instructions with human feedback.pdfproceedings.neurips.cc