Control protocols don’t always need to know which models are scheming — LessWrong
These are my personal views. • To detect if an agent is taking a catastrophically dangerous action, you might want to monitor its actions using the s…
x Control protocols don’t always need to know which models are scheming — LessWrong AI Frontpage 39 Control protocols don’t always need to know which models are scheming by Fabien Roger 26th Apr 2026 7 min read 1 39 These are my personal views. To detect if an agent is taking a catastrophically dangerous action, you might want to monitor its actions using the smartest model that is too weak to be a schemer. But knowing what models are weak enough that they are unlikely to scheme is difficult, which puts you in a difficult spot: take a model too strong and it might actually be a schemer and lie
Explore this link on the map →saved by
related reading
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- How to prevent collusion when using untrusted models to monitor each other — AI Alignment Forumalignmentforum.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Agentic Monitoring for AI Control — LessWronglesswrong.com
- 2312.06942arxiv.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Recent Redwood Research project proposals — AI Alignment Forumalignmentforum.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- Ctrl-Z: Controlling AI Agents via Resampling — LessWronglesswrong.com