Carl Guo on X: "New Paper 📄: LMs just want to explain themselves! When we SFT an LM on explanations of its own behaviors, do they learn to actually introspect, or do they merely imitate the original training distribution? We find evidence for the former. Despite training on a static set of https://t.co/AnznpunrJ6" / X
New Paper 📄: LMs just want to explain themselves! When we SFT an LM on explanations of its own behaviors, do they learn to actually introspect, or do they merely imitate the original training distribution? We find evidence for the former. Despite training on a static set of
Carl Guo @CarlGuo866 New Paper 📄: LMs just want to explain themselves! When we SFT an LM on explanations of its own behaviors, do they learn to actually introspect, or do they merely imitate the original training distribution? We find evidence for the former. Despite training on a static set of explanations from a base model, the SFT-ed model explains its own current behaviors better than the base model’s behaviors, tracking behavioral drift even when we don’t explicitly train it to. We call this introspective coupling: self-explanations track a model’s own behavior as that behavior changes,
Explore this link on the map →related reading
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- LLMs can learn about themselves by introspection — LessWronglesswrong.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- Transformer Circuits Threadtransformer-circuits.pub
- How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? — LessWronglesswrong.com
- [2602.02639] A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behaviorarxiv.org
- Position: It's Time to Optimize for Self-Consistencytime-for-consistency.github.io
- [2602.02639] A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behaviorarxiv.org