[2601.17087] Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
Abstract:Agentic benchmarks increasingly rely on LLM-simulated users to scalably evaluate agent performance, yet the robustness, validity, and fairness of this approach remain unexamined. Through a user study with participants across the United States, India, Kenya, and Nigeria, we investigate whether LLM-simulated users serve as reliable proxies for real human users in evaluating agents on {\tau}-Bench retail tasks. We find that user simulation lacks robustness, with agent success rates varying up to 9 percentage points across different user LLMs. Furthermore, evaluations using simulated users exhibit systematic miscalibration, underestimating agent performance on challenging tasks and overestimating it on moderately difficult ones. African American Vernacular English (AAVE) speakers experience consistently worse success rates and calibration errors than Standard American English (SAE) speakers, with disparities compounding significantly with age. We also find simulated users to be a differentially effective proxy for different populations, performing worst for AAVE and Indian English speakers. Additionally, simulated users introduce conversational artifacts and surface different failure patterns than human users. These findings demonstrate that current evaluation practices risk misrepresenting agent capabilities across diverse user populations and may obscure real-world deployment challenges.
View PDF HTML (experimental) Abstract:Agentic benchmarks increasingly rely on LLM-simulated users to scalably evaluate agent performance, yet the robustness, validity, and fairness of this approach remain unexamined. Through a user study with participants across the United States, India, Kenya, and Nigeria, we investigate whether LLM-simulated users serve as reliable proxies for real human users in evaluating agents on {\tau}-Bench retail tasks. We find that user simulation lacks robustness, with agent success rates varying up to 9 percentage points across different user LLMs. Furthermore,…
saved by
related reading
- Quantifying the Utility of User Simulators for Building Collaborative LLM Assistantsarxiv.org
- Simulating Users with State Alignment Beats Response Imitationhumanlm.stanford.edu
- LLMs Get Lost in Evolving User Intentarxiv.org
- Group | Sherry Tongshuang Wucs.cmu.edu
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Today, we are releasing a research preview of our user model, along with a set of evaluations designed to measure how faithfully user models capture human behavior.persimmon.humansand.ai
- Training and Evaluating User Language Modelsarxiv.org
- Building Effective AI Agents \ Anthropicanthropic.com
- HumanLM: Simulating Users with State Alignment Beats Response Imitationarxiv.org
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- [2304.03442] Generative Agents: Interactive Simulacra of Human Behaviorarxiv.org
- Building Effective AI Agents \ Anthropicanthropic.com