flâneur — a map of the web's best reading

How our new Control Red Team is stress-testing frontier monitors | AISI Work

aisi.gov.uk · 1,316 words · saved by 1 readers

Early learnings from red-teaming the internal monitors of frontier AI companies, and our perspectives on the open problems that remain.

As LLM agents become more autonomous, they have more opportunities to cause harm. To mitigate risk, frontier AI developers, such as OpenAI and Anthropic , are already deploying agents under the watch of a monitor: a separate LLM that reviews the agent’s actions and flags them if they are dangerous. The growing field of AI control research aims to design these kinds of safety measures, and evaluate whether they would be robust even if agents were trying to evade detection. For over two years, AISI’s Red Team has been evaluating misuse safeguards : protections against humans deliberately eliciti

Explore this link on the map →

related reading