Meta-level adversarial evaluation of oversight techniques might allow robust measurement of their adequacy — LessWrong
Summary: Many proposed AGI alignment procedures involve taking a pretrained model and training it using rewards from an oversight process to get a policy. These procedures might fail when the oversight procedure is locally inadequate: that is, if the model is able to trick the oversight process into giving good rewards for bad actions. In this post, we propose evaluating the local adequacy of oversight by constructing adversarial policies for oversight processes. Specifically, we propose constructing behaviors that a particular oversight process evaluates favorably but that we know to be bad via other means, such as additional held-out information or more expensive oversight processes. We think that this form of adversarial evaluation is a crucial part of ensuring that oversight processes are robust enough to oversee dangerously powerful models. A core element of many scenarios where AI ends up disempowering humanity (e.g. “Without specific countermeasures”) are oversight failures: tha
x Meta-level adversarial evaluation of oversight techniques might allow robust measurement of their adequacy — LessWrong Redwood Research AI Frontpage 101 Meta-level adversarial evaluation of oversight techniques might allow robust measurement of their adequacy by Buck , ryan_greenblatt 26th Jul 2023 AI Alignment Forum 1 min read 19 101 Ω 53 Summary: Many proposed AGI alignment procedures involve taking a pretrained model and training it using rewards from an oversight process to get a policy . These procedures might fail when the oversight procedure is locally inadequate: that is, if the mode
Explore this link on the map →related reading
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Reflections On The Feasibility Of Scalable-Oversight — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Research Areas in Benchmark Design and Evaluation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org