✳flâneur — a map of the web's best reading
How our new Control Red Team is stress-testing frontier monitors | AISI Work
aisi.gov.uk · 1,316 words · saved by 1 readers
Early learnings from red-teaming the internal monitors of frontier AI companies, and our perspectives on the open problems that remain.
As LLM agents become more autonomous, they have more opportunities to cause harm. To mitigate risk, frontier AI developers, such as OpenAI and Anthropic , are already deploying agents under the watch of a monitor: a separate LLM that reviews the agent’s actions and flags them if they are dangerous. The growing field of AI control research aims to design these kinds of safety measures, and evaluate whether they would be robust even if agents were trying to evade detection. For over two years, AISI’s Red Team has been evaluating misuse safeguards : protections against humans deliberately eliciti
Explore this link on the map →related reading
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- 2312.06942arxiv.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Recent Redwood Research project proposals — AI Alignment Forumalignmentforum.org
- Frontier Risk Report (February to March 2026) - METRmetr.org
- gdm-ai-control-roadmap.pdfstorage.googleapis.com
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- Agentic Monitoring for AI Control — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org