flâneur — a map of the web's best reading

Eliciting Language Model Behaviors with Investigator Agents | Transluce AI

transluce.org · 7,627 words · saved by 1 readers

Large language models (LMs) exhibit a wide range of behaviors due to their open-ended nature. As a result, it is hard to determine in advance what types of behavior a particular model can exhibit. For instance, even if a model fails to exhibit a certain harmful behavior across a set of test prompts, it's possible that a new jailbreak still elicits it [1, 2, 3]. Conversely, even if a model fails to exhibit a certain capability when prompted, it's possible that a better prompt (such as "Take a deep breath") will succeed [4, 5]. To address the open-ended input space of language models, we'd like tools that can search through this space to automatically surface specific behaviors of interest (e.g. instances of a specified failure mode). We call this task behavior elicitation. For example, eliciting “harmful responses” (commonly referred to as jailbreaking) enables us to identify safety vulnerabilities [6], and eliciting “hallucinations” helps us pinpoint knowledge gaps of LMs [7]. We train

Eliciting Language Model Behaviors with Investigator Agents Xiang Lisa Li * , Neil Chowdhury * , Daniel D. Johnson , Tatsunori Hashimoto , Percy Liang , Sarah Schwettmann , Jacob Steinhardt * Equal contribution. Correspondence to xlisali@stanford.edu, neil@transluce.org Transluce | Published: October 23, 2024 February 3, 2025: An updated version of this report is available on arXiv . Introduction Large language models (LMs) exhibit a wide range of behaviors due to their open-ended nature. As a result, it is hard to determine in advance what types of behavior a particular model can exhibit. For

Explore this link on the map →

related reading