Cost-Effective Constitutional Classifiers via Representation Re-use
Hoagy Cunningham, Alwin Peng, Jerry Wei, Euan Ong, Fabien Roger, Linda Petrini, Misha Wagner, Vladimir Mikulik, Mrinank Sharma TL;DR: We study cost-effective jailbreak detection. Instead of using a dedicated jailbreak classifier, we repurpose the computations that AI models already perform by fine-tuning just the final layer or using linear probes on intermediate activations. Our fine-tuned final layer detectors outperform standalone classifiers a quarter the size of the base model, while linear probes achieve performance comparable to a classifier 2% of the size of the policy model with virtually no additional computational cost. Probe classifiers also function effectively as first-stages in two-stage classification pipelines, further improving the cost-performance tradeoff. These methods could dramatically reduce the computational overhead of jailbreak detection, though further testing with adaptive adversarial attacks is needed. Sufficiently advanced AI systems could pose catastroph
Authors Hoagy Cunningham, Alwin Peng, Jerry Wei, Euan Ong, Fabien Roger, Linda Petrini, Misha Wagner, Vladimir Mikulik, Mrinank Sharma TL;DR: We study cost-effective jailbreak detection. Instead of using a dedicated jailbreak classifier, we repurpose the computations that AI models already perform by fine-tuning just the final layer or using linear probes on intermediate activations. Our fine-tuned final layer detectors outperform standalone classifiers a quarter the size of the base model, while linear probes achieve performance comparable to a classifier 2% of the size of the policy…
saved by
related reading
- Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaksarxiv.org
- [2602.15001] Boundary Point Jailbreaking of Black-Box LLMsarxiv.org
- [2510.09714] All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Languagearxiv.org
- Pre-deployment auditing can catch an overt saboteuralignment.anthropic.com
- [2512.11949] Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitorsarxiv.org
- [2501.18837] Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teamingarxiv.org
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Detecting and preventing distillation attacks \ Anthropicanthropic.com
- FAR.AI Leaderboard 2026leaderboard.far.ai
- Scaling Activation Oracles to Trillion-Parameter Modelstransluce.org
- Seeing in Pangram Space | Pangrampangram.com
- Goodfire AIgoodfire.ai