Automation collapse — LessWrong
If we validate automated alignment research through empirical testing, the safety assurance work will be similar to that needed for human-written alignment algorithms.
x Automation collapse — LessWrong AI-Assisted Alignment AI Frontpage 72 Automation collapse by Geoffrey Irving , Tomek Korbak , Benjamin Hilton 21st Oct 2024 AI Alignment Forum 9 min read 9 72 Ω 38 Summary: If we validate automated alignment research through empirical testing, the safety assurance work will still need to be done by humans, and will be similar to that needed for human-written alignment algorithms. Three levels of automated AI safety Automating AI safety means developing some algorithm which takes in data and outputs safe, highly-capable AI systems. Let’s imagine three ways of d
Explore this link on the map →saved by
related reading
- Prospects for Alignment Automation: Interpretability Case Study — LessWronglesswrong.com
- Automation collapse — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- Automated Alignment is Harder Than You Think — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- [2605.06390] Automated alignment is harder than you thinkarxiv.org
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com