flâneur

Cost-Effective Constitutional Classifiers via Representation Re-use

alignment.anthropic.com · 4,967 words · saved by 2 readers

Hoagy Cunningham, Alwin Peng, Jerry Wei, Euan Ong, Fabien Roger, Linda Petrini, Misha Wagner, Vladimir Mikulik, Mrinank Sharma TL;DR: We study cost-effective jailbreak detection. Instead of using a dedicated jailbreak classifier, we repurpose the computations that AI models already perform by fine-tuning just the final layer or using linear probes on intermediate activations. Our fine-tuned final layer detectors outperform standalone classifiers a quarter the size of the base model, while linear probes achieve performance comparable to a classifier 2% of the size of the policy model with virtually no additional computational cost. Probe classifiers also function effectively as first-stages in two-stage classification pipelines, further improving the cost-performance tradeoff. These methods could dramatically reduce the computational overhead of jailbreak detection, though further testing with adaptive adversarial attacks is needed. Sufficiently advanced AI systems could pose catastroph

Authors Hoagy Cunningham, Alwin Peng, Jerry Wei, Euan Ong, Fabien Roger, Linda Petrini, Misha Wagner, Vladimir Mikulik, Mrinank Sharma TL;DR: We study cost-effective jailbreak detection. Instead of using a dedicated jailbreak classifier, we repurpose the computations that AI models already perform by fine-tuning just the final layer or using linear probes on intermediate activations. Our fine-tuned final layer detectors outperform standalone classifiers a quarter the size of the base model, while linear probes achieve performance comparable to a classifier 2% of the size of the policy…

saved by

related reading