[2502.13329] Language Models Can Predict Their Own Behavior
Abstract:The text produced by language models (LMs) can exhibit specific `behaviors,' such as a failure to follow alignment training, that we hope to detect and react to during deployment. Identifying these behaviors can often only be done post facto, i.e., after the entire text of the output has been generated. We provide evidence that there are times when we can predict how an LM will behave early in computation, before even a single token is generated. We show that probes trained on the internal representation of input tokens alone can predict a wide range of eventual behaviors over the entire output sequence. Using methods from conformal prediction, we provide provable bounds on the estimation error of our probes, creating precise early warning systems for these behaviors. The conformal probes can identify instances that will trigger alignment failures (jailbreaking) and instruction-following failures, without requiring a single token to be generated. An early warning system built on the probes reduces jailbreaking by 91%. Our probes also show promise in pre-emptively estimating how confident the model will be in its response, a behavior that cannot be detected using the output text alone. Conformal probes can preemptively estimate the final prediction of an LM that uses Chain-of-Thought (CoT) prompting, hence accelerating inference. When applied to an LM that uses CoT to perform text classification, the probes drastically reduce inference costs (65% on average across 27 datasets), with negligible accuracy loss. Encouragingly, probes generalize to unseen datasets and perform better on larger models, suggesting applicability to the largest of models in real-world settings.
[2502.13329] Language Models Can Predict Their Own Behavior Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computation and Language arXiv:2502.13329 (cs) [Submitted on 18 Feb 2025 ( v1 ), last revised 22 Sep 2025 (this version, v2)] Title: Language Models Can Predict Their Own Behavior Authors: Dhananjay Ashok , Jonathan May View a PDF of the paper titled Language Models Can Predict Their Own Behavior, by Dhananjay Ashok and 1 other authors View PDF HTML (experimental) Abstract: T
Explore this link on the map →related reading
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- How confessions can keep language models honest | OpenAIopenai.com
- Surfacing Pathological Behaviors in Language Models | Transluce AItransluce.org
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2304.11082] Fundamental Limitations of Alignment in Large Language Modelsarxiv.org
- Eliciting Language Model Behaviors with Investigator Agents | Transluce AItransluce.org
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com