Can we safely automate alignment research? - Joe Carlsmith
joecarlsmith.com · 19,212 words · saved by 3 readers
It's really important; we have a real shot; there are a lot of ways we can fail.
Can we safely automate alignment research? - Joe Carlsmith How do we solve the alignment problem? / Part 6 Can we safely automate alignment research? Contents hide 1. Introduction 1.1 Executive summary 2. Why is automating alignment research so important? 3. Alignment MVPs 3.1 What if neither of these approaches are viable? 3.2 Alignment MVPs don’t imply “hand-off” 4. Why might automated alignment research fail? 5. Evaluation failures 5.1 Output-focused and process-focused evaluation 5.2 Human output-focused evaluation 5.3 Scalable oversight 5.4 Process-focused techniques 6 Comparisons with ot
saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Tips for Empirical Alignment Research — AI Alignment Forumalignmentforum.org
- The Universe from an Intentional Stancecasparoesterheld.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Automated Alignment is Harder Than You Think — LessWronglesswrong.com
- A minimal viable product for alignment - by Jan Leikealigned.substack.com
- Defining alignment research — LessWronglesswrong.com
- Readings on the nature of alignment researchcasparoesterheld.com
- Automated alignment is harder than you thinkarxiv.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- [2605.06390] Automated alignment is harder than you thinkarxiv.org