Discovering Language Model Behaviors with Model-Written Evaluations — LessWrong
“Discovering Language Model Behaviors with Model-Written Evaluations” is a new Anthropic paper by Ethan Perez et al. that I (Evan Hubinger) also coll…
x Discovering Language Model Behaviors with Model-Written Evaluations — LessWrong Language Models (LLMs) AI Frontpage 100 Discovering Language Model Behaviors with Model-Written Evaluations by evhub , Ethan Perez 20th Dec 2022 AI Alignment Forum 1 min read 34 100 Ω 43 This is a linkpost for https://www.anthropic.com/model-written-evals.pdf “ Discovering Language Model Behaviors with Model-Written Evaluations ” is a new Anthropic paper by Ethan Perez et al. that I (Evan Hubinger) also collaborated on. I think the results in this paper are quite interesting in terms of what they demonstrate abou
saved by
related reading
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- How confessions can keep language models honest | OpenAIopenai.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- [2501.11120] Tell me about yourself: LLMs are aware of their learned behaviorsarxiv.org
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Foundation Models for Oversight | Transluce AItransluce.org
- Self-CTRL: Self-Consistency Training with Reinforcement Learningarxiv.org
- [2212.09251] Discovering Language Model Behaviors with Model-Written Evaluationsarxiv.org
- Language Models Learn to Mislead Humans via RLHFarxiv.org
- Towards Understanding Sycophancy in Language Models — LessWronglesswrong.com
- Surfacing Pathological Behaviors in Language Models | Transluce AItransluce.org