✳flâneur — a map of the web's best reading
Automated Alignment is Harder Than You Think — LessWrong
lesswrong.com · 2,973 words · saved by 1 readers
Summary This is a summary of a paper published by the alignment team at UK AISI. Read the full paper here. …
x Automated Alignment is Harder Than You Think — LessWrong AI Frontpage 2026 Top Fifty: 14 % 143 Automated Alignment is Harder Than You Think by Aleksandr Bowkis , Marie_DB , Jacob Pfau , Geoffrey Irving 14th May 2026 Linkpost for arxiv.org 4 min read 7 143 Summary This is a summary of a paper published by the alignment team at UK AISI. Read the full paper here . AI research agents may help solve ASI alignment, for example via the following plan: Build agents that can do empirical alignment work (e.g.~writing code, running experiments, designing evaluations and red teaming) and confirm they ar
Explore this link on the map →saved by
related reading
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- [2605.06390] Automated alignment is harder than you thinkarxiv.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Automation collapse — LessWronglesswrong.com
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Why I’m optimistic about our alignment approachaligned.substack.com
- Automation collapse — AI Alignment Forumalignmentforum.org