[2605.06390] Automated alignment is harder than you think
Abstract:A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when research agents are not scheming to deliberately sabotage alignment work, this plan could produce compelling but catastrophically misleading safety assessments resulting in the unintentional deployment of misaligned AI. This could happen because alignment research involves many hard-to-supervise fuzzy tasks (tasks without clear evaluation criteria, for which human judgement is systematically flawed). Consequently, research outputs will contain systematic, undetected errors, and even correct outputs could be incorrectly aggregated into overconfident safety assessments. This problem is likely to be worse for automated alignment research than for human-generated alignment research for several reasons: 1) optimisation pressure means agent-generated mistakes are concentrated among those that human reviewers are least likely to catch; 2) agents are likely to produce errors that do not resemble human mistakes; 3) AI-generated alignment solutions may involve arguments humans cannot evaluate; and 4) shared weights, data and training processes may make AI outputs more correlated than human equivalents. Therefore, agents must be trained to reliably perform hard-to-supervise fuzzy tasks. Generalisation and scalable oversight are the leading candidates for achieving this but both face novel challenges in the context of automated alignment.
AUTOMATED A LIGNMENT IS H ARDER T HAN Y OU T HINK Aleksandr Bowkis Marie Davidsen Buhl Jacob Pfau AI Security Institute AI Security Institute AI Security Institute aleksandr.bowkis@dsit.gov.uk marie.buhl@dsit.gov.uk jacob.pfau@dsit.gov.uk Geoffrey…
related reading
- Automated alignment is harder than you thinkarxiv.org
- Automated Alignment is Harder Than You Think — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Jacob Pfau on Musings on the Alignment Problemaligned.substack.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org