Settings – ControlArena
control-arena.aisi.org.uk · 6,383 words · saved by 1 readers
Open-source framework for AI Control evaluations
Settings – ControlArena Source Overview A Setting in ControlArena defines a complete evaluation environment for testing control protocols. Settings simulate realistic risk scenarios where an AI agent performs legitimate work (the main task ) while potentially attempting malicious objectives (the side task ). For example, in the original AI Control paper , agents solved programming problems from the APPS dataset as their main task, while side tasks involved inserting subtle backdoors that pass tests but fail on specific secret inputs. Settings are implemented by extending control_arena.settings
related reading
- GitHub - UKGovernmentBEIS/control-arena: ControlArena is a collection of settings, model organisms and protocols - for running control experiments. · GitHubgithub.com
- Together AI | The AI Native Cloudtogether.ai
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Cookbookcookbook.openai.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- LangChain: the open agent platform to own your intelligencelangchain.com
- Building reliable AI agents · parth sareenparthsareen.com
- Goodfire AIgoodfire.ai
- The Agent Stackvercel.com
- [2604.15384] LinuxArena: A Control Setting for AI Agents in Live Production Software Environmentsarxiv.org
- Lakera – Test your AI hacking skillsgandalf.lakera.ai
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com