[2503.13657] Why Do Multi-Agent LLM Systems Fail?
Abstract:Despite enthusiasm for Multi-Agent LLM Systems (MAS), their performance gains on popular benchmarks are often minimal. This gap highlights a critical need for a principled understanding of why MAS fail. Addressing this question requires systematic identification and analysis of failure patterns. We introduce MAST-Data, a comprehensive dataset of 1600+ annotated traces collected across 7 popular MAS frameworks. MAST-Data is the first multi-agent system dataset to outline the failure dynamics in MAS for guiding the development of better future systems. To enable systematic classification of failures for MAST-Data, we build the first Multi-Agent System Failure Taxonomy (MAST). We develop MAST through rigorous analysis of 150 traces, guided closely by expert human annotators and validated by high inter-annotator agreement (kappa = 0.88). This process identifies 14 unique modes, clustered into 3 categories: (i) system design issues, (ii) inter-agent misalignment, and (iii) task verification. To enable scalable annotation, we develop an LLM-as-a-Judge pipeline with high agreement with human annotations. We leverage MAST and MAST-Data to analyze failure patterns across models (GPT4, Claude 3, Qwen2.5, CodeLlama) and tasks (coding, math, general agent), demonstrating improvement headrooms from better MAS design. Our analysis provides insights revealing that identified failures require more sophisticated solutions, highlighting a clear roadmap for future research. We publicly release our comprehensive dataset (MAST-Data), the MAST, and our LLM annotator to facilitate widespread research and development in MAS.
Abstract:Despite enthusiasm for Multi-Agent LLM Systems (MAS), their performance gains on popular benchmarks are often minimal. This gap highlights a critical need for a principled understanding of why MAS fail. Addressing this question requires systematic identification and analysis of failure patterns. We introduce MAST-Data, a comprehensive dataset of 1600+ annotated traces collected across 7 popular MAS frameworks. MAST-Data is the first multi-agent system dataset to outline the failure dynamics in MAS for guiding the development of better future systems. To enable systematic classificatio
Explore this link on the map →related reading
- Getting Up to Speed on Multi-Agent Systems, Part 1: The Landscapechristophermeiklejohn.com
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- How we built our multi-agent research system \ Anthropicanthropic.com
- Building Effective AI Agents \ Anthropicanthropic.com
- Don't Sleep on Single-agent Systems | Sep 26, 2024all-hands.dev
- Building Effective AI Agents \ Anthropicanthropic.com
- Don’t Build Multi-Agents | Cognitioncognition.ai
- Agent Evaluation: A Detailed Guidecameronrwolfe.substack.com
- Theory of Mind for Multi-Agent Collaboration via Large Language Modelsarxiv.org
- Scalable Multi-Robot Collaboration with Large Language Models: Centralized or Decentralized Systems?arxiv.org
- Agent Observability and Tracingarize.com
- What are agents? 🤔 · Hugging Facehuggingface.co