Automatically Jailbreaking Frontier Language Models with Investigator Agents | Transluce AI
We train investigator agents using reinforcement learning to generate natural language jailbreaks for 48 high-risk tasks involving CBRN materials, explosives, and illegal drugs. Our results show success against models including GPT-5-main (78%), Claude Sonnet 4 (92%), and Gemini 2.5 Pro (90%). We find that small open-weight investigator models can successfully attack frontier target models, demonstrating an approach to cost-effective red-teaming.
Automatically Jailbreaking Frontier Language Models with Investigator Agents Neil Chowdhury* , Sarah Schwettmann , Jacob Steinhardt * Correspondence to: neil@transluce.org Transluce | Published: September 3, 2025 We train investigator agents using reinforcement learning to generate natural language jailbreaks for 48 high-risk tasks involving CBRN materials, explosives, and illegal drugs. Our results show success against models including GPT-5-main (78%), Claude Sonnet 4 (92%), and Gemini 2.5 Pro (90%). We find that small open-weight investigator models can successfully attack frontier target m
related reading
- gpt-4.pdfcdn.openai.com
- FAR.AI Leaderboard 2026leaderboard.far.ai
- [2602.15001] Boundary Point Jailbreaking of Black-Box LLMsarxiv.org
- Lakera – Test your AI hacking skillsgandalf.lakera.ai
- Eliciting Language Model Behaviors with Investigator Agents | Transluce AItransluce.org
- Frontier Risk Report (February to March 2026) - METRmetr.org
- PostTrainBenchposttrainbench.com
- [2603.02202] Frontier Models Can Take Actions at Low Probabilitiesarxiv.org
- Cost-Effective Constitutional Classifiers via Representation Re-usealignment.anthropic.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Modular Pretraining Enables Access Controlalignment.anthropic.com
- BenchmarkList: Track the Frontier of AI Capabilitiesbenchmarklist.com