flâneur — a map of the web's best reading

Agentic Monitoring for AI Control — LessWrong

lesswrong.com · 4,221 words · saved by 1 readers

tl;dr - We present an introduction and initial investigation into agentic monitoring for AI control protocols. We explain the motivation behind giving monitors opportunities to be agentic, demonstrate and discuss some early experimental results, and recommend directions for future work, highlighting possible risks. AI control aims to limit catastrophic consequences from using advanced AI systems that might be scheming against us. One of the most fundamental control measures we could employ is trusted monitoring - prompting a weaker trusted model to monitor an untrusted model’s outputs and ‘raise the alarm’ if it notices anything suspicious. One avenue to improve this line of defence is to make our trusted monitors better. Current approaches tend to apply monitoring directly, either to the untrusted model’s actions or reasoning trace (or both). We hope that there are sufficient clues here to catch unwanted behaviour before it’s executed: blatantly malicious actions, or reasoning steps l

x Agentic Monitoring for AI Control — LessWrong AI Control AI Frontpage 10 Agentic Monitoring for AI Control by LAThomson 27th Oct 2025 11 min read 0 10 tl;dr - We present an introduction and initial investigation into agentic monitoring for AI control protocols. We explain the motivation behind giving monitors opportunities to be agentic, demonstrate and discuss some early experimental results, and recommend directions for future work, highlighting possible risks. Introduction AI control aims to limit catastrophic consequences from using advanced AI systems that might be scheming against us.

Explore this link on the map →

related reading