Expanding AI Control from Models to Harnesses — LessWrong
Most AI control research such as LinuxArena and Ctrl-Z only gives the red team basic agents which only have access to tools. Yet in 2026, usage of AI…
x Expanding AI Control from Models to Harnesses — LessWrong AI "Agent" Scaffolds AI Control AI Frontpage 17 Expanding AI Control from Models to Harnesses by fastfedora 15th Jul 2026 24 min read 0 17 Most AI control research such as LinuxArena and Ctrl-Z only gives the red team basic agents which only have access to tools. Yet in 2026, usage of AI within frontier labs has moved to agent harnesses that have access to skills, memory, subagents, external services, compaction and more. At the same time, Claude Code and Codex have both implemented their own version of both action-based and source co
related reading
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Reading Listblog.redwoodresearch.org
- gdm-ai-control-roadmap.pdfstorage.googleapis.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- How our new Control Red Team is stress-testing frontier monitors | AISI Workaisi.gov.uk
- An overview of areas of control workblog.redwoodresearch.org
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workaisi.gov.uk
- [2607.07368] Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitorsarxiv.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- Agentic Monitoring for AI Control — LessWronglesswrong.com