flâneur — a map of the web's best reading

Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forum

alignmentforum.org · 4,340 words · saved by 3 readers

In this post, I’ll present a research direction that I’m interested in for alignment of pretrained language models. TL;DR: Force a language model to think out loud, and use the reasoning itself as a channel for oversight. If this agenda is successful, it could defeat deception, power-seeking, and other forms of disapproved reasoning. This direction is broadly actionable now. In recent publications, prompting pretrained language models to work through logical reasoning problems step-by-step has provided a boost to their capabilities. I claim that this externalized reasoning process may be used for alignment if three conditions are met: If these conditions hold, we should be able to detect and avoid models that reason through convergent instrumental goals to be deceptive, power-seeking, non-myopic, or reason through other processes of which we don’t approve. Reasoning oversight should provide stronger guarantees of alignment than oversight on model outputs alone, since we would get insig

x Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forum Chain-of-Thought Alignment Inner Alignment Language Models (LLMs) MATS Program Outer Alignment AI Frontpage 57 Externalized reasoning oversight: a research direction for language model alignment by tamera 3rd Aug 2022 8 min read 23 57 Summary In this post, I’ll present a research direction that I’m interested in for alignment of pretrained language models. TL;DR: Force a language model to think out loud, and use the reasoning itself as a channel for oversight. If this agenda is successful

Explore this link on the map →

saved by

related reading