Reflections On The Feasibility Of Scalable-Oversight — LessWrong
This post was inspired by discussions on the feasibility of scaling RLHF for aligning AGI during EAG 2023 in Oakland. The first draft was finished du…
x Reflections On The Feasibility Of Scalable-Oversight — LessWrong RLHF AI Frontpage 11 Reflections On The Feasibility Of Scalable-Oversight by Felix Hofstätter 10th Mar 2023 15 min read 0 11 This post was inspired by discussions on the feasibility of scaling RLHF for aligning AGI during EAG 2023 in Oakland. The first draft was finished during the Apart Research Thinkaton the following week. Thanks to Lee Sharkey, Daniel Ziegler, Tom Henighan, Stephen Casper, Lucius Bushnaq, Niclas Brand, Stefan Hemersheim for the valuable conversations I had during those events. The conclusions and inevitable
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Measuring progress on scalable oversightanthropic.com
- Can we scale human feedback for complex AI tasks? An intro to scalable oversight.aisafetyfundamentals.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com