✳flâneur — a map of the web's best reading
Can we safely automate alignment research? - Joe Carlsmith
joecarlsmith.com · 19,212 words · saved by 2 readers
It's really important; we have a real shot; there are a lot of ways we can fail.
Can we safely automate alignment research? - Joe Carlsmith How do we solve the alignment problem? / Part 6 Can we safely automate alignment research? Contents hide 1. Introduction 1.1 Executive summary 2. Why is automating alignment research so important? 3. Alignment MVPs 3.1 What if neither of these approaches are viable? 3.2 Alignment MVPs don’t imply “hand-off” 4. Why might automated alignment research fail? 5. Evaluation failures 5.1 Output-focused and process-focused evaluation 5.2 Human output-focused evaluation 5.3 Scalable oversight 5.4 Process-focused techniques 6 Comparisons with ot
Explore this link on the map →saved by
related reading
- Tips for Empirical Alignment Research — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- I Would Have Solved Alignment, But I Was Worried That Would Advance Timelines — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Automated Alignment is Harder Than You Think — LessWronglesswrong.com
- A minimal viable product for alignment - by Jan Leikealigned.substack.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org
- [2605.06390] Automated alignment is harder than you thinkarxiv.org
- Why I’m optimistic about our alignment approachaligned.substack.com
- Why I’m optimistic about our alignment approachaligned.substack.com