[2602.02639] A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
Abstract:LLM self-explanations are often presented as a promising tool for AI oversight, yet their faithfulness to the model's true reasoning process is poorly understood. Existing faithfulness metrics have critical limitations, typically relying on identifying unfaithfulness via adversarial prompting or detecting reasoning errors. These methods overlook the predictive value of explanations. We introduce Normalized Simulatability Gain (NSG), a general and scalable metric based on the idea that a faithful explanation should allow an observer to learn a model's decision-making criteria, and thus better predict its behavior on related inputs. We evaluate 18 frontier proprietary and open-weight models, e.g., Gemini 3, GPT-5.2, and Claude 4.5, on 7,000 counterfactuals from popular datasets covering health, business, and ethics. We find self-explanations substantially improve prediction of model behavior (11-37% NSG). Self-explanations also provide more predictive information than explanations generated by external models, even when those models are stronger. This implies an advantage from self-knowledge that external explanation methods cannot replicate. Our approach also reveals that, across models, 5-15% of self-explanations are egregiously misleading. Despite their imperfections, we show a positive case for self-explanations: they encode information that helps predict model behavior.
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior Harry Mayne * 1 Justin Singh Kang * 2 Dewi Gould 3 Kannan Ramchandran 2 Adam Mahdi 1 Noah Y. Siegel 4 5 Abstract Reference patient Counterfactual patient LLM self-explanations are often presented as a Sex꞉ M…
related reading
- [2602.02639] A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behaviorarxiv.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Self-CTRL: Self-Consistency Training with Reinforcement Learningarxiv.org
- [2511.08579] Training Language Models to Explain Their Own Computationsarxiv.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- Position: It's Time to Optimize for Self-Consistencytime-for-consistency.github.io
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Training Language Models to Explain Their Own Computationsarxiv.org
- Training Language Models to Explain Their Own Computationsarxiv.org
- How confessions can keep language models honest | OpenAIopenai.com
- [2505.17120] Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisionsarxiv.org