Guardian Angels: LLM Personalization for Productivity and Security · Gwern.net
I propose an approach for highly personalized LLMs, for near-future productivity gains and personal info/cybersecurity against increasingly powerful LLMs: they should, in the spirit of uploading, try to emulate the user’s values and preferences in order to amplify the principal—not replace them. I discuss a package of techniques and proposals to accomplish such ‘guardian angels’; dynamic evaluation of LLMs combined with active learning and elicitation and heavy inner-monologue search/data-augmentation.
--- title: "Guardian Angels: LLM Personalization for Productivity and Security" description: "I propose an approach for highly personalized LLMs, for near-future productivity gains and personal info/cybersecurity against increasingly powerful LLMs: they should, in the spirit of uploading, try to emulate the user’s values and preferences in order to amplify the principal—not replace them. I discuss a package of techniques and proposals to accomplish such ‘guardian angels’; dynamic evaluation of LLMs combined with active learning and elicitation and heavy inner-monologue search/data-augmentation
Explore this link on the map →saved by
- Justin Wang
- Uzay Girit
- Lydia Nottingham
- Vincent Cheng
- Timothy Kostolansky
- Abhay Sheshadri
- Kushal Thaman
- Vincent Huang
- Jackson Mowatt Gok
- Clem von Stengel
- Arunim Agarwal
related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- [2606.06614] Re-Centering Humans in LLM Personalizationarxiv.org
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- AI in 2025: gestalt — LessWronglesswrong.com
- 2025 LLM Year in Review – karpathykarpathy.bearblog.dev
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- The persona selection model — LessWronglesswrong.com
- GenAI Handbookgenai-handbook.github.io