Agentic Monitoring for AI Control — LessWrong
tl;dr - We present an introduction and initial investigation into agentic monitoring for AI control protocols. We explain the motivation behind giving monitors opportunities to be agentic, demonstrate and discuss some early experimental results, and recommend directions for future work, highlighting possible risks. AI control aims to limit catastrophic consequences from using advanced AI systems that might be scheming against us. One of the most fundamental control measures we could employ is trusted monitoring - prompting a weaker trusted model to monitor an untrusted model’s outputs and ‘raise the alarm’ if it notices anything suspicious. One avenue to improve this line of defence is to make our trusted monitors better. Current approaches tend to apply monitoring directly, either to the untrusted model’s actions or reasoning trace (or both). We hope that there are sufficient clues here to catch unwanted behaviour before it’s executed: blatantly malicious actions, or reasoning steps l
x Agentic Monitoring for AI Control — LessWrong AI Control AI Frontpage 10 Agentic Monitoring for AI Control by LAThomson 27th Oct 2025 11 min read 0 10 tl;dr - We present an introduction and initial investigation into agentic monitoring for AI control protocols. We explain the motivation behind giving monitors opportunities to be agentic, demonstrate and discuss some early experimental results, and recommend directions for future work, highlighting possible risks. Introduction AI control aims to limit catastrophic consequences from using advanced AI systems that might be scheming against us.
Explore this link on the map →related reading
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- 2312.06942arxiv.org
- Control protocols don’t always need to know which models are scheming — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Early work on monitorability evaluations - METRmetr.org
- How to prevent collusion when using untrusted models to monitor each other — AI Alignment Forumalignmentforum.org
- Recent Redwood Research project proposals — AI Alignment Forumalignmentforum.org
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org