flâneur — a map of the web's best reading

When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors

arxiv.org · 38,492 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.

\uselogo \correspondingauthor semmons@google.com \reportnumber 186324 When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors Scott Emmons \thepa Erik Jenner \thepa David K. Elson \thepa Rif A. Saurous Google, Paradigms of Intelligence Team Senthooran Rajamanoharan \thepa Heng Chen \thepa Irhum Shafkat \thepa Rohin Shah \thepa Abstract While chain-of-thought (CoT) monitoring is an appealing AI safety defense, recent work on “unfaithfulness” has cast doubt on its reliability. These findings highlight an important failure mode, particularly when CoT acts as a post-hoc rati

Explore this link on the map →

saved by

related reading