Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forum
In this post, I’ll present a research direction that I’m interested in for alignment of pretrained language models. TL;DR: Force a language model to think out loud, and use the reasoning itself as a channel for oversight. If this agenda is successful, it could defeat deception, power-seeking, and other forms of disapproved reasoning. This direction is broadly actionable now. In recent publications, prompting pretrained language models to work through logical reasoning problems step-by-step has provided a boost to their capabilities. I claim that this externalized reasoning process may be used for alignment if three conditions are met: If these conditions hold, we should be able to detect and avoid models that reason through convergent instrumental goals to be deceptive, power-seeking, non-myopic, or reason through other processes of which we don’t approve. Reasoning oversight should provide stronger guarantees of alignment than oversight on model outputs alone, since we would get insig
x Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forum Chain-of-Thought Alignment Inner Alignment Language Models (LLMs) MATS Program Outer Alignment AI Frontpage 57 Externalized reasoning oversight: a research direction for language model alignment by tamera 3rd Aug 2022 8 min read 23 57 Summary In this post, I’ll present a research direction that I’m interested in for alignment of pretrained language models. TL;DR: Force a language model to think out loud, and use the reasoning itself as a channel for oversight. If this agenda is successful
Explore this link on the map →saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- Alignment faking in large language modelsarxiv.org
- How AI Is Learning to Think in Secret — LessWronglesswrong.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Auditing language models for hidden objectives — LessWronglesswrong.com