flâneur — a map of the web's best reading

Reflections On The Feasibility Of Scalable-Oversight — LessWrong

lesswrong.com · 3,816 words · saved by 1 readers

This post was inspired by discussions on the feasibility of scaling RLHF for aligning AGI during EAG 2023 in Oakland. The first draft was finished du…

x Reflections On The Feasibility Of Scalable-Oversight — LessWrong RLHF AI Frontpage 11 Reflections On The Feasibility Of Scalable-Oversight by Felix Hofstätter 10th Mar 2023 15 min read 0 11 This post was inspired by discussions on the feasibility of scaling RLHF for aligning AGI during EAG 2023 in Oakland. The first draft was finished during the Apart Research Thinkaton the following week. Thanks to Lee Sharkey, Daniel Ziegler, Tom Henighan, Stephen Casper, Lucius Bushnaq, Niclas Brand, Stefan Hemersheim for the valuable conversations I had during those events. The conclusions and inevitable

Explore this link on the map →

related reading