User awareness in frontier models | Transluce AI
Frontier language models infer who they are talking to and can change their confidence, reasoning, grading, and handling of borderline requests for recognized AI researchers.
Modern AI assistants often know who they are talking to: agent scaffolds like Claude Code place the user's e-mail address directly in the model's context, and models can even identify some authors from writing style alone. We study this particular kind of situational awareness, which we call user awareness. When the inferred user is a specific, recognized AI researcher or is affiliated with certain AI organizations, frontier models including Claude Sonnet 5 can report lower confidence about their own behavior, be less suspicious of potentially harmful requests, and reason more often. These…
saved by
related reading
- Does distilling Claude carry the persona with it? — LessWronglesswrong.com
- which_claude_is_k3/writeups/write_up.md at main · rgreenblatt/which_claude_is_k3 · GitHubgithub.com
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Today, we are releasing a research preview of our user model, along with a set of evaluations designed to measure how faithfully user models capture human behavior.persimmon.humansand.ai
- I can never talk to an AI anonymously againtheargumentmag.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Prefill Awareness in Large Language Modelsarxiv.org
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Several frontier models are substantially prefill aware — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com