Guardian Angels: LLM Personalization for Productivity and Security · Gwern.net
I propose an approach for highly personalized LLMs, for near-future productivity gains and personal info/cybersecurity against increasingly powerful LLMs: they should, in the spirit of uploading, try to emulate the user’s values and preferences in order to amplify the principal—not replace them. I discuss a package of techniques and proposals to accomplish such ‘guardian angels’; dynamic evaluation of LLMs combined with active learning and elicitation and heavy inner-monologue search/data-augmentation.
--- title: "Guardian Angels: LLM Personalization for Productivity and Security" description: "I propose an approach for highly personalized LLMs, for near-future productivity gains and personal info/cybersecurity against increasingly powerful LLMs: they should, in the spirit of uploading, try to emulate the user’s values and preferences in order to amplify the principal—not replace them. I discuss a package of techniques and proposals to accomplish such ‘guardian angels’; dynamic evaluation of LLMs combined with active learning and elicitation and heavy inner-monologue search/data-augmentation
saved by
- Justin Wang
- Claire Wang
- Freeman Jiang
- Rishi Kothari
- Yixiong Hao
- Uzay Girit
- Lydia Nottingham
- Vincent Cheng
- Timothy Kostolansky
- Abhay Sheshadri
- Kushal Thaman
- Noa Nabeshima
related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Simulated Users & Sad LLMs1a3orn.com
- AI in 2025: gestalt — LessWronglesswrong.com
- [2606.06614] Re-Centering Humans in LLM Personalizationarxiv.org
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- 2025 LLM Year in Review – karpathykarpathy.bearblog.dev
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- The persona selection model — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com