flâneur — a map of the web's best reading

Meta-level adversarial evaluation of oversight techniques might allow robust measurement of their adequacy — LessWrong

lesswrong.com · 7,175 words · saved by 1 readers

Summary: Many proposed AGI alignment procedures involve taking a pretrained model and training it using rewards from an oversight process to get a policy. These procedures might fail when the oversight procedure is locally inadequate: that is, if the model is able to trick the oversight process into giving good rewards for bad actions. In this post, we propose evaluating the local adequacy of oversight by constructing adversarial policies for oversight processes. Specifically, we propose constructing behaviors that a particular oversight process evaluates favorably but that we know to be bad via other means, such as additional held-out information or more expensive oversight processes. We think that this form of adversarial evaluation is a crucial part of ensuring that oversight processes are robust enough to oversee dangerously powerful models. A core element of many scenarios where AI ends up disempowering humanity (e.g. “Without specific countermeasures”) are oversight failures: tha

x Meta-level adversarial evaluation of oversight techniques might allow robust measurement of their adequacy — LessWrong Redwood Research AI Frontpage 101 Meta-level adversarial evaluation of oversight techniques might allow robust measurement of their adequacy by Buck , ryan_greenblatt 26th Jul 2023 AI Alignment Forum 1 min read 19 101 Ω 53 Summary: Many proposed AGI alignment procedures involve taking a pretrained model and training it using rewards from an oversight process to get a policy . These procedures might fail when the oversight procedure is locally inadequate: that is, if the mode

Explore this link on the map →

related reading