[2607.20734] LLMs Get Lost in Evolving User Intent
Abstract:As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, revising, and reshaping it as the conversation unfolds. Despite this, LLMs are still predominantly evaluated or trained in single-turn, fully-specified settings, leaving open a fundamental question: how well do LLMs track and act on user intent as it evolves over the course of a conversation? To study this, we introduce a framework that transforms static, single-turn tasks into dynamic multi-turn conversations in which the user's intent evolves across turns--incrementally revealed, revised, and at times redirected mid-conversation--while preserving each task's original evaluation protocol, enabling existing benchmarks to be reused as controlled testbeds without new annotation. Across multiple tasks, we surface a consistent phenomenon: strong static-setting performance does not transfer to the evolving-intent setting, with substantial drops across model families. Our findings point to a fundamental gap: today's LLMs do not yet faithfully track and act on the user's evolving intent, a capability invisible to static evaluation yet critical for future collaborative agents.
View PDF HTML (experimental) Abstract:As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, revising, and reshaping it as the conversation unfolds. Despite this, LLMs are still predominantly evaluated or trained in single-turn, fully-specified settings, leaving open a fundamental question: how well do LLMs track and act on user intent as it evolves over the course of a conversation? To study…
saved by
related reading
- Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluationsarxiv.org
- [2508.19227] Generative Interfaces for Language Modelsarxiv.org
- user_interactions.pdfself-distillation.github.io
- Quantifying the Utility of User Simulators for Building Collaborative LLM Assistantsarxiv.org
- Behavior Latticing: Inferring User Motivations from Unstructured Interactionsarxiv.org
- PhD_thesis_Shirley_Wu_final.pdfcs.stanford.edu
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- [2602.03429] DiscoverLLM: From Executing Intents to Discovering Themarxiv.org
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Building Effective AI Agents \ Anthropicanthropic.com
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- Patterns for Building LLM-based Systems & Productseugeneyan.com