Automation collapse — AI Alignment Forum
Summary: If we validate automated alignment research through empirical testing, the safety assurance work will still need to be done by humans, and will be similar to that needed for human-written alignment algorithms. Automating AI safety means developing some algorithm which takes in data and outputs safe, highly-capable AI systems. Let’s imagine three ways of developing this algorithm: There is a level (4) of automated safety, where you trust the AIs helping with automation to know when they are right, and thus do not need to check the results. This corresponds to the “deference” level in Clymer et al. We treat this as out of scope for this doc, as it requires a level of capability beyond where automated alignment plans normally kick in. We will argue that: In detail, level (2) – AI-written algorithm, checked empirically – consists of This picture is quite general: the research environment could support reinforcement learning of agents with custom scalable oversight schemes, process
x Automation collapse — AI Alignment Forum AI-Assisted Alignment AI Frontpage 38 Automation collapse by Geoffrey Irving , Tomek Korbak , Benjamin Hilton 21st Oct 2024 9 min read 9 38 Summary: If we validate automated alignment research through empirical testing, the safety assurance work will still need to be done by humans, and will be similar to that needed for human-written alignment algorithms. Three levels of automated AI safety Automating AI safety means developing some algorithm which takes in data and outputs safe, highly-capable AI systems. Let’s imagine three ways of developing this
Explore this link on the map →related reading
- Automation collapse — LessWronglesswrong.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- Automated Alignment is Harder Than You Think — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- [2605.06390] Automated alignment is harder than you thinkarxiv.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Prospects for Alignment Automation: Interpretability Case Study — LessWronglesswrong.com
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org