Automated Alignment is Harder Than You Think — LessWrong
lesswrong.com · 2,973 words · saved by 2 readers
Summary This is a summary of a paper published by the alignment team at UK AISI. Read the full paper here. …
x Automated Alignment is Harder Than You Think — LessWrong AI Frontpage 2026 Top Fifty: 14 % 143 Automated Alignment is Harder Than You Think by Aleksandr Bowkis , Marie_DB , Jacob Pfau , Geoffrey Irving 14th May 2026 Linkpost for arxiv.org 4 min read 7 143 Summary This is a summary of a paper published by the alignment team at UK AISI. Read the full paper here . AI research agents may help solve ASI alignment, for example via the following plan: Build agents that can do empirical alignment work (e.g.~writing code, running experiments, designing evaluations and red teaming) and confirm they ar
saved by
related reading
- Automation collapse — LessWronglesswrong.com
- Automated alignment is harder than you thinkarxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- [2605.06390] Automated alignment is harder than you thinkarxiv.org
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Jacob Pfau on Musings on the Alignment Problemaligned.substack.com
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org