Reflections On The Feasibility Of Scalable-Oversight — LessWrong
This post was inspired by discussions on the feasibility of scaling RLHF for aligning AGI during EAG 2023 in Oakland. The first draft was finished du…
x Reflections On The Feasibility Of Scalable-Oversight — LessWrong RLHF AI Frontpage 11 Reflections On The Feasibility Of Scalable-Oversight by Felix Hofstätter 10th Mar 2023 15 min read 0 11 This post was inspired by discussions on the feasibility of scaling RLHF for aligning AGI during EAG 2023 in Oakland. The first draft was finished during the Apart Research Thinkaton the following week. Thanks to Lee Sharkey, Daniel Ziegler, Tom Henighan, Stephen Casper, Lucius Bushnaq, Niclas Brand, Stefan Hemersheim for the valuable conversations I had during those events. The conclusions and inevitable
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Meta-level adversarial evaluation of oversight techniques might allow robust measurement of their adequacy — LessWronglesswrong.com