Expanding AI Control from Models to Harnesses — LessWrong
Most AI control research such as LinuxArena and Ctrl-Z only gives the red team basic agents which only have access to tools. Yet in 2026, usage of AI…
x Expanding AI Control from Models to Harnesses — LessWrong AI "Agent" Scaffolds AI Control AI Frontpage 17 Expanding AI Control from Models to Harnesses by fastfedora 15th Jul 2026 24 min read 0 17 Most AI control research such as LinuxArena and Ctrl-Z only gives the red team basic agents which only have access to tools. Yet in 2026, usage of AI within frontier labs has moved to agent harnesses that have access to skills, memory, subagents, external services, compaction and more. At the same time, Claude Code and Codex have both implemented their own version of both action-based and source co
Explore this link on the map →related reading
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- gdm-ai-control-roadmap.pdfstorage.googleapis.com
- Scaling Managed Agents: Decoupling the brain from the hands \ Anthropicanthropic.com
- Agentic Monitoring for AI Control — LessWronglesswrong.com
- A basic systems architecture for AI agents that do autonomous research — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- Arjun Virkarjunvirk.com