FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users
Effective personalization of LLMs is critical for a broad range of user-interfacing applications such as virtual assistants and content curation. Inspired by the strong in-context capabilities of LLMs, we propose few-shot preference optimization (FSPO), an algorithm for LLM personalization that reframes reward modeling as a meta-learning problem. Under FSPO, an LLM learns to quickly infer a personalized reward function for a user via a few labeled preferences. FSPO also utilizes user description rationalization (RAT) to encourage better reward modeling and instruction following, recovering performance with the oracle user description. Since real-world preference data is challenging to collect at scale, we propose careful design choices to construct synthetic preference datasets for personalization, generating over 1M synthetic personalized preferences using publicly available LLMs. To successfully transfer from synthetic data to real users, we find it crucial for the data to exhibit bo
\reportnumber \correspondingauthor anikait@stanford.edu. Project Website: https://fewshot-preference-optimization.github.io/ FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users Anikait Singh Stanford University Sheryl Hsu Stanford University Kyle Hsu Stanford University Eric Mitchell Stanford University OpenAI Stefano Ermon Stanford University Tatsunori Hashimoto Stanford University Archit Sharma Stanford University Google DeepMind Chelsea Finn Stanford University Abstract Effective personalization of LLMs is critical for a broad range of user-interfacing applicatio
related reading
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- [2606.06614] Re-Centering Humans in LLM Personalizationarxiv.org
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Probing Persona-Dependent Preferences in Language Modelsarxiv.org
- Fine-tune Llama 2 with DPOhuggingface.co
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- [2302.08582] Pretraining Language Models with Human Preferencesarxiv.org
- The persona selection model — LessWronglesswrong.com
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io