[2601.04603] Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
Abstract:We introduce enhanced Constitutional Classifiers that deliver production-grade jailbreak robustness with dramatically reduced computational costs and refusal rates compared to previous-generation defenses. Our system combines several key insights. First, we develop exchange classifiers that evaluate model responses in their full conversational context, which addresses vulnerabilities in last-generation systems that examine outputs in isolation. Second, we implement a two-stage classifier cascade where lightweight classifiers screen all traffic and escalate only suspicious exchanges to more expensive classifiers. Third, we train efficient linear probe classifiers and ensemble them with external classifiers to simultaneously improve robustness and reduce computational costs. Together, these techniques yield a production-grade system achieving a 40x computational cost reduction compared to our baseline exchange classifier, while maintaining a 0.05% refusal rate on production traffic. Through extensive red-teaming comprising over 1,700 hours, we demonstrate strong protection against universal jailbreaks -- no attack on this system successfully elicited responses to all eight target queries comparable in detail to an undefended model. Our work establishes Constitutional Classifiers as practical and efficient safeguards for large language models.
Authors:Hoagy Cunningham, Jerry Wei, Zihan Wang, Andrew Persic, Alwin Peng, Jordan Abderrachid, Raj Agarwal, Bobby Chen, Austin Cohen, Andy Dau, Alek Dimitriev, Rob Gilson, Logan Howard, Yijin Hua, Jared Kaplan, Jan Leike, Mu Lin, Christopher Liu, Vladimir Mikulik, Rohit Mittapalli, Clare O'Hara, Jin Pan, Nikhil Saxena, Alex Silverstein, Yue Song, Xunjie Yu, Giulio Zhou, Ethan Perez, Mrinank Sharma View PDF HTML (experimental) Abstract:We introduce enhanced Constitutional Classifiers that deliver production-grade jailbreak robustness with dramatically reduced computational costs and…
saved by
related reading
- Cost-Effective Constitutional Classifiers via Representation Re-usealignment.anthropic.com
- [2602.15001] Boundary Point Jailbreaking of Black-Box LLMsarxiv.org
- [2501.18837] Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teamingarxiv.org
- What is input/output filtering in AI safety?blog.bluedot.org
- [2510.09714] All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Languagearxiv.org
- Modular Pretraining Enables Access Controlalignment.anthropic.com
- FAR.AI Leaderboard 2026leaderboard.far.ai
- Anthropic February 2026 Risk Report: SecureBio's External Reviewsecurebio.org
- Detecting and preventing distillation attacks \ Anthropicanthropic.com
- Nicholas Carlininicholas.carlini.com
- Security incident disclosure — July 2026huggingface.co
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com