[2603.03303] HumanLM: Simulating Users with State Alignment Beats Response Imitation
Abstract:Large Language Models (LLMs) are increasingly used to simulate how specific users respond to a given context, enabling more user-centric applications that rely on user feedback. However, existing user simulators mostly imitate surface-level patterns and language styles, which fail to reflect the underlying states of real users (e.g., beliefs and emotions). To address these limitations, we propose a novel training framework, HumanLM, which builds user simulators that accurately reflect real users. Our key insight is that, in addition to generating responses, the model should generate natural-language latent states that align with ground-truth responses through reinforcement learning. These latent states correspond to a set of psychologically grounded state dimensions that drive how real users respond. HumanLM further synthesizes these aligned latent states into responses that accurately represent real users. For extensive evaluation, we develop Humanual, a comprehensive benchmark for simulating real users based on public data. Humanual consists of six large-scale datasets with 26k users and 216k responses in total, spanning diverse tasks such as generating user responses to daily life issues, political blogs, and chat sessions with LLM assistants. Across datasets, HumanLM significantly outperforms alternative approaches, achieving an average relative improvement of 16.3% in alignment scores from an LLM judge. In a real-time simulation study with 111 participants, HumanLM achieves the highest similarity to real user responses and competitive human-likeness scores.
View PDF HTML (experimental) Abstract:Large Language Models (LLMs) are increasingly used to simulate how specific users respond to a given context, enabling more user-centric applications that rely on user feedback. However, existing user simulators mostly imitate surface-level patterns and language styles, which fail to reflect the underlying states of real users (e.g., beliefs and emotions). To address these limitations, we propose a novel training framework, HumanLM, which builds user simulators that accurately reflect real users. Our key insight is that, in addition to generating…
saved by
related reading
- Training and Evaluating User Language Modelsarxiv.org
- Simulating Users with State Alignment Beats Response Imitationhumanlm.stanford.edu
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Quantifying the Utility of User Simulators for Building Collaborative LLM Assistantsarxiv.org
- Today, we are releasing a research preview of our user model, along with a set of evaluations designed to measure how faithfully user models capture human behavior.persimmon.humansand.ai
- Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluationsarxiv.org
- [2606.06614] Re-Centering Humans in LLM Personalizationarxiv.org
- Simulated Users & Sad LLMs1a3orn.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- [2305.14387] AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedbackarxiv.org
- [2304.03442] Generative Agents: Interactive Simulacra of Human Behaviorarxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com