flâneur — a map of the web's best reading

Neural Chameleons: LLMs Can Learn to Evade Activation Monitors

neuralchameleons.com · 137 words · saved by 1 readers

Language models can learn to evade activation monitors with 0-shot transfer to unseen deception and harmfulness detectors, while retaining capabilities.

Neural Chameleons Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors Authors: Max McGuinness (MATS), Alex Serrano (MATS, UC Berkeley), Luke Bailey (Stanford University), Scott Emmons (MATS) Summary We show that language models can learn a general skill to evade activation monitors, with 0-shot transfer to unseen deception and harmfulness monitors from the literature, with minimal effect on model capabilities. We call these "Neural Chameleons". Method We used a "stress-testing" approach: fine-tuning models to evade simple probes for benign concepts when given a tri

Explore this link on the map →

related reading