[2608.21664] Measuring Activation Control in Large Language Models
Abstract:Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especially when evaluation-aware models exhibit scheming or deception. However, if models can also control their own activations, deception could extend into the latent space itself. With this in mind, we introduce the Activation Controllability Benchmark to quantify the extent to which models can modulate their residual stream via natural-language instruction. Across model families and capability levels, we find that most LLMs can control the direction and magnitude of their residual stream activations with some degree of temporal resolution, though performance varies considerably across models. In simple tasks, this level of control can evade activation-based monitoring methods (including linear probes, natural language autoencoders, activation oracles, and the Jacobian lens), albeit imperfectly. These results suggest that control over the activation space itself could become a confound for monitoring as introspective capabilities increase; therefore, we recommend that frontier labs and evaluators track activation controllability in future models.
Measuring Activation Control in Large Language Models Marek Mateusz Kowalski∗ , Joshua Fonseca Rivera, Uzay Macar, David Demitri Africa† Abstract make faithful self-explanation a native model capability (Li arXiv:2608.21664v1 [cs.AI] 21 Aug 2026 Safe deployment of increasingly capable models will likely 2026; Guo et al. 2026), but write access is also a risk. A…
saved by
related reading
- [2512.11949] Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitorsarxiv.org
- [2606.04071] Covert Influence Between Language Modelsarxiv.org
- Inside a Neural Chameleonjacksonmowattgok.com
- [2607.14111] Introspection Fine-Tuning (IFT): Training Small LLMs to Introspectarxiv.org
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- Scaling Activation Oracles to Trillion-Parameter Modelstransluce.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org
- Neural Chameleons: LLMs Can Learn to Evade Activation Monitorsneuralchameleons.com
- [2508.00161] Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMsarxiv.org