Why I’m optimistic about our alignment approach
OpenAI’s approach to alignment research involves perfecting RLHF, AI-assisted human evaluation, and automated alignment research. Why is this a good strategy? What are the reasons to be optimistic about it? My optimism stems from five sources: Positive updates about AI. A lot of developments over the last few years have made AI systems more favorable to alignment than they looked initially, both in terms of how the AI tech tree is shaking out and the empirical evidence on alignment we’ve gathered so far. A more modest goal. We’re not trying to solve all alignment problems. We’re just trying to align a system that’s capable enough to make more alignment progress than we can. Evaluation is easier than generation. This is a very general principle that holds across many domains. It’s true for alignment research as well. We’re setting ourselves up for iteration. We can set ourselves up for iterative, measurable improvements on our alignment path. Conviction in language models. Language mode
Why I’m optimistic about our alignment approach Some arguments in favor and responses to common objections Jan Leike Dec 05, 2022 62 16 2 Share OpenAI’s approach to alignment research involves perfecting RLHF , AI-assisted human evaluation , and automated alignment research . Why is this a good strategy? What are the reasons to be optimistic about it? My optimism stems from five sources: Positive updates about AI. A lot of developments over the last few years have made AI systems more favorable to alignment than they looked initially, both in terms of how the AI tech tree is shaking out and th
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Why I’m optimistic about our alignment approachaligned.substack.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Tips for Empirical Alignment Research — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- A minimal viable product for alignment - by Jan Leikealigned.substack.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Automated Alignment is Harder Than You Think — LessWronglesswrong.com