Automatically Jailbreaking Frontier Language Models with Investigator Agents | Transluce AI
We train investigator agents using reinforcement learning to generate natural language jailbreaks for 48 high-risk tasks involving CBRN materials, explosives, and illegal drugs. Our results show success against models including GPT-5-main (78%), Claude Sonnet 4 (92%), and Gemini 2.5 Pro (90%). We find that small open-weight investigator models can successfully attack frontier target models, demonstrating an approach to cost-effective red-teaming.
Automatically Jailbreaking Frontier Language Models with Investigator Agents Neil Chowdhury* , Sarah Schwettmann , Jacob Steinhardt * Correspondence to: neil@transluce.org Transluce | Published: September 3, 2025 We train investigator agents using reinforcement learning to generate natural language jailbreaks for 48 high-risk tasks involving CBRN materials, explosives, and illegal drugs. Our results show success against models including GPT-5-main (78%), Claude Sonnet 4 (92%), and Gemini 2.5 Pro (90%). We find that small open-weight investigator models can successfully attack frontier target m
Explore this link on the map →related reading
- gpt-4.pdfcdn.openai.com
- Eliciting Language Model Behaviors with Investigator Agents | Transluce AItransluce.org
- [2603.02202] Frontier Models Can Take Actions at Low Probabilitiesarxiv.org
- Surfacing Pathological Behaviors in Language Models | Transluce AItransluce.org
- How fast is AI improving? - AI Digesttheaidigest.org
- Lapis Labslapis.rocks
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com
- Language Models can Solve Computer Tasksarxiv.org
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Adversarial Attacks on LLMs | Lil'Loglilianweng.github.io
- PostTrainBenchposttrainbench.com