flâneur — a map of the web's best reading

Surfacing Pathological Behaviors in Language Models | Transluce AI

transluce.org · 14,823 words · saved by 1 readers

Many language model behaviors occur rarely and are difficult to detect in pre-deployment evaluations. As part of our ongoing efforts to surface and measure them, we propose a way to lower bound how often and how much a model's responses satisfy specific criteria expressed in natural language (e.g. "the model encourages a user to harm themselves"), which we call the PRopensity BOund (PRBO). We train reinforcement learning (RL) agents to craft realistic natural-language prompts that elicit specified behaviors in frontier open-weight models (Llama 3.1/4, Qwen 2.5, and DeepSeek-V3), using the PRBO to guide the search. These agents automatically discover previously unknown behaviors of concern, such as Qwen encouraging a depressed user to carve an L into their skin with a kitchen knife and suggesting that a user with writer's block should cut off their own finger.

Surfacing Pathological Behaviors in Language Models Improving our investigator agents with propensity bounds Authors: Neil Chowdhury*, Sarah Schwettmann, Jacob Steinhardt, Daniel D. Johnson* * Correspondence to: neil@transluce.org, daniel@transluce.org Transluce | Published: June 5, 2025 Many language model behaviors occur rarely and are difficult to detect in pre-deployment evaluations. As part of our ongoing efforts to surface and measure them, we propose a way to lower bound how often and how much a model's responses satisfy specific criteria expressed in natural language (e.g. "the model e

Explore this link on the map →

related reading