✳flâneur — a map of the web's best reading
Why I’m optimistic about our alignment approach
aligned.substack.com · 6,588 words · saved by 1 readers
Some arguments in favor and responses to common objections
Why I’m optimistic about our alignment approach Some arguments in favor and responses to common objections Jan Leike Dec 05, 2022 62 16 2 Share OpenAI’s approach to alignment research involves perfecting RLHF , AI-assisted human evaluation , and automated alignment research . Why is this a good strategy? What are the reasons to be optimistic about it? My optimism stems from five sources: Positive updates about AI. A lot of developments over the last few years have made AI systems more favorable to alignment than they looked initially, both in terms of how the AI tech tree is shaking out and th
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Why I’m optimistic about our alignment approachaligned.substack.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- Tips for Empirical Alignment Research — AI Alignment Forumalignmentforum.org
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- A minimal viable product for alignment - by Jan Leikealigned.substack.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org