[2501.11120] Tell me about yourself: LLMs are aware of their learned behaviors
Abstract:We study behavioral self-awareness -- an LLM's ability to articulate its behaviors without requiring in-context examples. We finetune LLMs on datasets that exhibit particular behaviors, such as (a) making high-risk economic decisions, and (b) outputting insecure code. Despite the datasets containing no explicit descriptions of the associated behavior, the finetuned LLMs can explicitly describe it. For example, a model trained to output insecure code says, ``The code I write is insecure.'' Indeed, models show behavioral self-awareness for a range of behaviors and for diverse evaluations. Note that while we finetune models to exhibit behaviors like writing insecure code, we do not finetune them to articulate their own behaviors -- models do this without any special training or examples. Behavioral self-awareness is relevant for AI safety, as models could use it to proactively disclose problematic behaviors. In particular, we study backdoor policies, where models exhibit unexpected behaviors only under certain trigger conditions. We find that models can sometimes identify whether or not they have a backdoor, even without its trigger being present. However, models are not able to directly output their trigger by default. Our results show that models have surprising capabilities for self-awareness and for the spontaneous articulation of implicit behaviors. Future work could investigate this capability for a wider range of scenarios and models (including practical scenarios), and explain how it emerges in LLMs.
[2501.11120] Tell me about yourself: LLMs are aware of their learned behaviors --> Computer Science > Computation and Language arXiv:2501.11120 (cs) [Submitted on 19 Jan 2025] Title: Tell me about yourself: LLMs are aware of their learned behaviors Authors: Jan Betley , Xuchan Bao , Martín Soto , Anna Sztyber-Betley , James Chua , Owain Evans View a PDF of the paper titled Tell me about yourself: LLMs are aware of their learned behaviors, by Jan Betley and 5 other authors View PDF HTML (experimental) Abstract: We study behavioral self-awareness -- an LLM's ability to articulate its behaviors w
saved by
related reading
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Self-CTRL: Self-Consistency Training with Reinforcement Learningarxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- Emergent introspective awareness in large language models \ Anthropicanthropic.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com
- User awareness in frontier modelstransluce.org
- [2505.17120] Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisionsarxiv.org