[2605.06390] Automated alignment is harder than you think
Abstract:A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when research agents are not scheming to deliberately sabotage alignment work, this plan could produce compelling but catastrophically misleading safety assessments resulting in the unintentional deployment of misaligned AI. This could happen because alignment research involves many hard-to-supervise fuzzy tasks (tasks without clear evaluation criteria, for which human judgement is systematically flawed). Consequently, research outputs will contain systematic, undetected errors, and even correct outputs could be incorrectly aggregated into overconfident safety assessments. This problem is likely to be worse for automated alignment research than for human-generated alignment research for several reasons: 1) optimisation pressure means agent-generated mistakes are concentrated among those that human reviewers are least likely to catch; 2) agents are likely to produce errors that do not resemble human mistakes; 3) AI-generated alignment solutions may involve arguments humans cannot evaluate; and 4) shared weights, data and training processes may make AI outputs more correlated than human equivalents. Therefore, agents must be trained to reliably perform hard-to-supervise fuzzy tasks. Generalisation and scalable oversight are the leading candidates for achieving this but both face novel challenges in the context of automated alignment.
Abstract:A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when research agents are not scheming to deliberately sabotage alignment work, this plan could produce compelling but catastrophically misleading safety assessments resulting in the unintentional deployment of misaligned AI. This could happen because alignment research involves many hard-to-supervise fuzzy tasks (tasks without clear evaluation criteria, for which human judgement is systematically fl
Explore this link on the map →related reading
- Automated Alignment is Harder Than You Think — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- ALIGNMENT - by vincent huang - a slice of my mindmindslice.substack.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Automation collapse — LessWronglesswrong.com
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- The hot mess theory of AI misalignment: More intelligent agents behave less coherently | Jascha’s blogsohl-dickstein.github.io