flâneur — a map of the web's best reading

Should We Train Against (CoT) Monitors? — LessWrong

lesswrong.com · 14,943 words · saved by 3 readers

The question I actually try to answer in this post is a broader one (that doesn't work as well as a title): Should we incorporate proxies for desired…

x Should We Train Against (CoT) Monitors? — LessWrong Aether AI Frontpage 50 Should We Train Against (CoT) Monitors? by RohanS 23rd Apr 2026 39 min read 7 50 The question I actually try to answer in this post is a broader one (that doesn't work as well as a title): Should we incorporate proxies for desired behavior into LLM alignment training? Epistemic status: My best guess. I tentatively claim that we should be more open to incorporating proxies for desired behavior into LLM training, but I try to clarify the spectrum of possible answers beyond just 'yes' and 'no,' and I try to present and s

Explore this link on the map →

saved by

related reading