flâneur — a map of the web's best reading

Automatically Jailbreaking Frontier Language Models with Investigator Agents | Transluce AI

transluce.org · 4,431 words · saved by 1 readers

We train investigator agents using reinforcement learning to generate natural language jailbreaks for 48 high-risk tasks involving CBRN materials, explosives, and illegal drugs. Our results show success against models including GPT-5-main (78%), Claude Sonnet 4 (92%), and Gemini 2.5 Pro (90%). We find that small open-weight investigator models can successfully attack frontier target models, demonstrating an approach to cost-effective red-teaming.

Automatically Jailbreaking Frontier Language Models with Investigator Agents Neil Chowdhury* , Sarah Schwettmann , Jacob Steinhardt * Correspondence to: neil@transluce.org Transluce | Published: September 3, 2025 We train investigator agents using reinforcement learning to generate natural language jailbreaks for 48 high-risk tasks involving CBRN materials, explosives, and illegal drugs. Our results show success against models including GPT-5-main (78%), Claude Sonnet 4 (92%), and Gemini 2.5 Pro (90%). We find that small open-weight investigator models can successfully attack frontier target m

Explore this link on the map →

related reading