[2501.18837] Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Abstract:Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes that require many model interactions, like manufacturing illegal substances at scale. To defend against these attacks, we introduce Constitutional Classifiers: safeguards trained on synthetic data, generated by prompting LLMs with natural language rules (i.e., a constitution) specifying permitted and restricted content. In over 3,000 estimated hours of red teaming, no red teamer found a universal jailbreak that could extract information from an early classifier-guarded LLM at a similar level of detail to an unguarded model across most target queries. On automated evaluations, enhanced classifiers demonstrated robust defense against held-out domain-specific jailbreaks. These classifiers also maintain deployment viability, with an absolute 0.38% increase in production-traffic refusals and a 23.7% inference overhead. Our work demonstrates that defending against universal jailbreaks while maintaining practical deployment viability is tractable.
[2501.18837] Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computation and Language arXiv:2501.18837 (cs) [Submitted on 31 Jan 2025] Title: Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming Authors: Mrinank Sharma , Meg Tong , Jesse Mu , Jerry Wei , Jorrit Kruthoff , Scott Goodfriend , Euan Ong , Alwin Peng , Raj Aga
Explore this link on the map →related reading
- Security incident disclosure — July 2026huggingface.co
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Nicholas Carlininicholas.carlini.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Claude Fable 5 and Claude Mythos 5 \ Anthropicanthropic.com
- Universal and Transferable Attacks on Aligned Language Modelsllm-attacks.org
- 2312.06942arxiv.org
- Automatically Jailbreaking Frontier Language Models with Investigator Agents | Transluce AItransluce.org
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- Redeploying Claude Fable 5 \ Anthropicanthropic.com
- Adversarial Attacks on LLMs | Lil'Loglilianweng.github.io
- How well do models follow their constitutions? — LessWronglesswrong.com