Evaluating LLMs for Impersonation
arxiv.org · 6,892 words · saved by 1 readers
N/A
Preprint IMPersona: Evaluating Individual Level LM Impersonation Quan Shi, Carlos E. Jimenez, Stephen Dong, Brian Seo, Caden Yao, Adam Kelch, Karthik Narasimhan Princeton Language and Intelligence (PLI), Princeton University {qbshi, carlosej}@princeton.edu Abstract arXiv:2504.04332v2…
related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Prompt Injection as Role Confusionrole-confusion.github.io
- The persona selection model — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Role-playing vs Self-modelling — LessWronglesswrong.com
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2606.06614] Re-Centering Humans in LLM Personalizationarxiv.org
- HumanLM: Simulating Users with State Alignment Beats Response Imitationarxiv.org
- The case for more ambitious language model evals — LessWronglesswrong.com