Introspection Adapters: Training LLMs to Report Their Learned Behaviors
When model developers or users fine-tune an LLM, this can induce behaviors that are unexpected, deliberately harmful, or hard to detect. It would be far easier to audit LLMs if they could simply describe their behaviors in natural language. Here, we study a scalable approach to rapidly identify learned behaviors of many LLMs derived from a shared base LLM. Given a model 𝑀 , our method works by finetuning models 𝑀 𝑖 from 𝑀 with implanted behaviors 𝑏 𝑖 ; the ( 𝑀 𝑖 , 𝑏 𝑖 ) pairs serve as labeled training data. We then train an introspection adapter (IA): a single LoRA adapter jointly trained across the finetunes 𝑀 𝑖 to cause them to verbalize their implanted behaviors. We find that this IA induces self-description of learned behaviors even in finetunes of 𝑀 that were trained in very different ways from the 𝑀 𝑖 . For example, IAs generalize to AuditBench, achieving state-of-the-art at identifying explicitly hidden concerning behaviors. IAs can also be used to de
Introspection Adapters: Training LLMs to Report Their Learned Behaviors Keshav Shenoy Li Yang Abhay Sheshadri Sören Mindermann Jack Lindsey Sam Marks Rowan Wang Abstract When model developers or users fine-tune an LLM, this can induce behaviors that are unexpected, deliberately harmful, or hard to detect. It would be far easier to audit LLMs if they could simply describe their behaviors in natural language. Here, we study a scalable approach to rapidly identify learned behaviors of many LLMs derived from a shared base LLM. Given a model M M , our method works by finetuning models M i M_{i} fro
Explore this link on the map →saved by
related reading
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2501.11120] Tell me about yourself: LLMs are aware of their learned behaviorsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2503.10965] Auditing language models for hidden objectivesarxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- How confessions can keep language models honest | OpenAIopenai.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- Auditing language models for hidden objectives — LessWronglesswrong.com
- AuditBenchalignment.anthropic.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org