flâneur

Inside a Neural Chameleon

jacksonmowattgok.com · 4,289 words · saved by 3 readers

Reverse-engineering how a language model hides from activation monitors

A language model that learned to hide# Neural Chameleons showed that models can learn a general skill for altering their neural activity to hide from completely new, unseen activation monitors.[1]Mainly tested using unseen linear probes but evasion also carried to some probe ensembles, MLPs, and attention monitors. How is this even possible? This seems like a bizarre meta-skill to be capable of. McGuinness et al. describes that: “When triggered, the model pushes its activations into a narrow, low-dimensional subspace, effectively steering them away from the regions a probe would flag as…

saved by

related reading