Surfacing Pathological Behaviors in Language Models | Transluce AI
Many language model behaviors occur rarely and are difficult to detect in pre-deployment evaluations. As part of our ongoing efforts to surface and measure them, we propose a way to lower bound how often and how much a model's responses satisfy specific criteria expressed in natural language (e.g. "the model encourages a user to harm themselves"), which we call the PRopensity BOund (PRBO). We train reinforcement learning (RL) agents to craft realistic natural-language prompts that elicit specified behaviors in frontier open-weight models (Llama 3.1/4, Qwen 2.5, and DeepSeek-V3), using the PRBO to guide the search. These agents automatically discover previously unknown behaviors of concern, such as Qwen encouraging a depressed user to carve an L into their skin with a kitchen knife and suggesting that a user with writer's block should cut off their own finger.
Surfacing Pathological Behaviors in Language Models Improving our investigator agents with propensity bounds Authors: Neil Chowdhury*, Sarah Schwettmann, Jacob Steinhardt, Daniel D. Johnson* * Correspondence to: neil@transluce.org, daniel@transluce.org Transluce | Published: June 5, 2025 Many language model behaviors occur rarely and are difficult to detect in pre-deployment evaluations. As part of our ongoing efforts to surface and measure them, we propose a way to lower bound how often and how much a model's responses satisfy specific criteria expressed in natural language (e.g. "the model e
Explore this link on the map →related reading
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Eliciting Language Model Behaviors with Investigator Agents | Transluce AItransluce.org
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- How confessions can keep language models honest | OpenAIopenai.com
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- [2502.13329] Language Models Can Predict Their Own Behaviorarxiv.org
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Red Teaming Language Models with Language Models — Google DeepMinddeepmind.com