[2604.15384] LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
Abstract:We introduce LinuxArena, a control setting in which agents operate directly on live, multi-service production environments. LinuxArena contains 20 environments, 1,671 main tasks representing legitimate software engineering work, and 184 side tasks representing safety failures such as data exfiltration and backdooring, making it the largest and most diverse control setting for software engineering to date. We validate LinuxArena is useful for control research by running sabotage evaluations, which measure whether attackers can complete side tasks while working on main tasks, and monitor evaluations, which measure a monitor model's ability to detect sabotage attempts. Against a GPT-5-nano trusted monitor at a 1\% step-wise false positive rate, Claude Opus 4.6 achieves roughly a 23% undetected sabotage success rate. We additionally release LaStraj, a dataset of human-crafted attack trajectories that evade monitors at substantially higher rates than any model-generated attacks we elicited, showing that current attack policies do not saturate LinuxArena. These results suggest that LinuxArena has meaningful headroom for both attackers and defenders, making it a strong testbed for developing and evaluating future control protocols.
LinuxArena: A Control Setting for AI Agents in Live Production Software Environments Redwood Research Tyler Tracy, Ram Potham, Nick Kuhn, Myles Heller, Anshul Khandelwal, Cody Rushing EquiStamp Henri Lemoine, Miguel Brandão, Tomáš Turlik, Adam Hanson, Josh Hills, Amy Ngo, Ram…
saved by
related reading
- [2607.07368] Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitorsarxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- GitHub - UKGovernmentBEIS/control-arena: ControlArena is a collection of settings, model organisms and protocols - for running control experiments. · GitHubgithub.com
- How our new Control Red Team is stress-testing frontier monitors | AISI Workaisi.gov.uk
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workaisi.gov.uk
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- BashArena and Control Setting Designblog.redwoodresearch.org
- 2312.06942arxiv.org
- Why it's hard to make settings for high-stakes control researchblog.redwoodresearch.org
- gdm-ai-control-roadmap.pdfstorage.googleapis.com
- An overview of areas of control workblog.redwoodresearch.org