flâneur

evhub's Shortform — LessWrong

lesswrong.com · saved by 1 readers

Comment by evhub - An intuition pump I really like for thinking about scalable oversight techniques is to think about how humans manage to produce scientific and philosophical progress. The goal of scalable oversight at a high level is to figure out how to incentivize models to output truth even when no human knows or can verify that truth. But that's exactly what the history of science is: society consistently figuring out how to converge towards truths that they didn't previously know. So how do humans do that, and are there lessons for how we should train models to do it in a similar way? 1. Humans use outcome signals. There are real, concrete problems in the world and some ideas actually solve those problems—in ways we can check and verify—and some ideas do not. Similarly, we can run actual empirical experiments, and some ideas correctly predict the results of those experiments, and some ideas do not. In the ML analogy, this corresponds to outcome-based RL. This is an obvious one, and clearly a very important one for humans. That being said, clearly this is not the only way that humans are able to produce progress: many scientific fields are able to continue making progress even without new experimental results, and in some cases like philosophy, without ever interfacing with experiment at all. Furthermore, and perhaps most importantly, outcome-based signals incentivize humans to make progress, but not necessarily to share that progress with the world—indeed, in many cases people just interested in outcome-based progress (like AI labs!) discover lots of truths but then decline to share them with others. 2. Humans use peer review. Perhaps a standard answer in many academic fields is that progress is incentivized by peer review and other mechanisms via which people in the field judge each others' work and assign status accordingly. In the ML analogy, this corresponds to optimization against a preference/grader model. I would also put debate in this category, si

saved by