✳flâneur — a map of the web's best reading
Settings – ControlArena
control-arena.aisi.org.uk · 6,383 words · saved by 1 readers
Open-source framework for AI Control evaluations
Settings – ControlArena Source Overview A Setting in ControlArena defines a complete evaluation environment for testing control protocols. Settings simulate realistic risk scenarios where an AI agent performs legitimate work (the main task ) while potentially attempting malicious objectives (the side task ). For example, in the original AI Control paper , agents solved programming problems from the APPS dataset as their main task, while side tasks involved inserting subtle backdoors that pass tests but fail on specific secret inputs. Settings are implemented by extending control_arena.settings
Explore this link on the map →related reading
- GitHub - UKGovernmentBEIS/control-arena: ControlArena is a collection of settings, model organisms and protocols - for running control experiments. · GitHubgithub.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Cookbookcookbook.openai.com
- Building reliable AI agents · parth sareenparthsareen.com
- Claude Code Cheat Sheetcc.storyfox.cz
- Arjun Virkarjunvirk.com
- Quantifying infrastructure noise in agentic coding evals \ Anthropicanthropic.com
- GitHub - microsoft/WindowsAgentArena: Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents. · GitHubgithub.com
- Gemini Enterprise Agent Platform (formerly Vertex AI) | Google Cloudcloud.google.com
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- AINews | AINewsnews.smol.ai
- API Overview | OpenAI API Referenceplatform.openai.com