Flipping the Dialogue: Training and Evaluating User Language Models
Conversations with LMs involve two participants: a human user leading the conversation, and an LM assistant responding to the user’s request. To satisfy this specific role, LMs are post-trained to be helpful assistants – optimized to produce exhaustive and well-structured responses, free of ambiguity and grammar errors. User utterances, on the other hand, are rarely perfected, with each user phrasing requests in unique ways, sometimes putting in partial effort at each turn and refining on the fly. To evaluate LM performance in realistic settings, prior work simulated users in multi-turn conversations, often prompting an LLM originally trained to be a helpful assistant to act as a user. However, we show that assistant LMs make for poor user simulators, with the surprising finding that better assistants yield worse simulators. Instead, we introduce purpose-built User Language Models (User LMs) - models post-trained to simulate human users in multi-turn conversations. Through various eval
Flipping the Dialogue: Training and Evaluating User Language Models Tarek Naous ††thanks: Work done while interning at Microsoft Research. Affiliation: Microsoft Research Affiliation: Georgia Institute of Technologytareknaous@gatech.edu ; plaban@microsoft.com Wei Xu Affiliation: Georgia Institute of Technologytareknaous@gatech.edu ; plaban@microsoft.com Jennifer Neville Affiliation: Microsoft Research Abstract Conversations with LMs involve two participants: a human user leading the conversation, and an LM assistant responding to the user’s request. To satisfy this specific role,…
saved by
related reading
- HumanLM: Simulating Users with State Alignment Beats Response Imitationarxiv.org
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Quantifying the Utility of User Simulators for Building Collaborative LLM Assistantsarxiv.org
- Today, we are releasing a research preview of our user model, along with a set of evaluations designed to measure how faithfully user models capture human behavior.persimmon.humansand.ai
- Simulating Users with State Alignment Beats Response Imitationhumanlm.stanford.edu
- Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluationsarxiv.org
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- The persona selection model — LessWronglesswrong.com
- Crossing the uncanny valley of conversational voice | Sesamesesame.com
- LLMs Get Lost in Evolving User Intentarxiv.org
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com