flâneur — a map of the web's best reading

Discovering Language Model Behaviors with Model-Written Evaluations — LessWrong

lesswrong.com · 7,361 words · saved by 2 readers

“Discovering Language Model Behaviors with Model-Written Evaluations” is a new Anthropic paper by Ethan Perez et al. that I (Evan Hubinger) also coll…

x Discovering Language Model Behaviors with Model-Written Evaluations — LessWrong Language Models (LLMs) AI Frontpage 100 Discovering Language Model Behaviors with Model-Written Evaluations by evhub , Ethan Perez 20th Dec 2022 AI Alignment Forum 1 min read 34 100 Ω 43 This is a linkpost for https://www.anthropic.com/model-written-evals.pdf “ Discovering Language Model Behaviors with Model-Written Evaluations ” is a new Anthropic paper by Ethan Perez et al. that I (Evan Hubinger) also collaborated on. I think the results in this paper are quite interesting in terms of what they demonstrate abou

Explore this link on the map →

saved by

related reading