Eliciting Language Model Behaviors with Investigator Agents | Transluce AI
Large language models (LMs) exhibit a wide range of behaviors due to their open-ended nature. As a result, it is hard to determine in advance what types of behavior a particular model can exhibit. For instance, even if a model fails to exhibit a certain harmful behavior across a set of test prompts, it's possible that a new jailbreak still elicits it [1, 2, 3]. Conversely, even if a model fails to exhibit a certain capability when prompted, it's possible that a better prompt (such as "Take a deep breath") will succeed [4, 5]. To address the open-ended input space of language models, we'd like tools that can search through this space to automatically surface specific behaviors of interest (e.g. instances of a specified failure mode). We call this task behavior elicitation. For example, eliciting “harmful responses” (commonly referred to as jailbreaking) enables us to identify safety vulnerabilities [6], and eliciting “hallucinations” helps us pinpoint knowledge gaps of LMs [7]. We train
Eliciting Language Model Behaviors with Investigator Agents Xiang Lisa Li * , Neil Chowdhury * , Daniel D. Johnson , Tatsunori Hashimoto , Percy Liang , Sarah Schwettmann , Jacob Steinhardt * Equal contribution. Correspondence to xlisali@stanford.edu, neil@transluce.org Transluce | Published: October 23, 2024 February 3, 2025: An updated version of this report is available on arXiv . Introduction Large language models (LMs) exhibit a wide range of behaviors due to their open-ended nature. As a result, it is hard to determine in advance what types of behavior a particular model can exhibit. For
Explore this link on the map →related reading
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Surfacing Pathological Behaviors in Language Models | Transluce AItransluce.org
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- How confessions can keep language models honest | OpenAIopenai.com
- Unsupervised Elicitationalignment.anthropic.com
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- 3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitationalignment.anthropic.com
- Mechanistically Eliciting Latent Behaviors in Language Models — LessWronglesswrong.com
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org