Discovering Language Model Behaviors with Model-Written Evaluations — LessWrong
“Discovering Language Model Behaviors with Model-Written Evaluations” is a new Anthropic paper by Ethan Perez et al. that I (Evan Hubinger) also coll…
x Discovering Language Model Behaviors with Model-Written Evaluations — LessWrong Language Models (LLMs) AI Frontpage 100 Discovering Language Model Behaviors with Model-Written Evaluations by evhub , Ethan Perez 20th Dec 2022 AI Alignment Forum 1 min read 34 100 Ω 43 This is a linkpost for https://www.anthropic.com/model-written-evals.pdf “ Discovering Language Model Behaviors with Model-Written Evaluations ” is a new Anthropic paper by Ethan Perez et al. that I (Evan Hubinger) also collaborated on. I think the results in this paper are quite interesting in terms of what they demonstrate abou
Explore this link on the map →saved by
related reading
- [2501.11120] Tell me about yourself: LLMs are aware of their learned behaviorsarxiv.org
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- [2212.09251] Discovering Language Model Behaviors with Model-Written Evaluationsarxiv.org
- Towards Understanding Sycophancy in Language Models — LessWronglesswrong.com
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- Surfacing Pathological Behaviors in Language Models | Transluce AItransluce.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- The case for more ambitious language model evals — LessWronglesswrong.com
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io