Understanding the recent criticism of the Chatbot Arena
The Chatbot Arena has become the go-to place for vibes-based evaluation of LLMs over the past two years. The project, originating at UC Berkeley, is home to a large community …
Understanding the recent criticism of the Chatbot Arena Simon Willison’s Weblog Subscribe Sponsored by: Atlassian - Give your agents a plan. Not a prompt. New Jira capabilities unlock full-context for AI-native software development. Assign tasks to Claude, Cursor, or GitHub Copilot, now directly from Jira. Learn more Understanding the recent criticism of the Chatbot Arena 30th April 2025 The Chatbot Arena has become the go-to place for vibes-based evaluation of LLMs over the past two years. The project, originating at UC Berkeley, is home to a large community of model enthusiasts who submit pr
Explore this link on the map →related reading
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- 2025: The year in LLMssimonwillison.net
- The bitter lesson of LLM evalsparsed.com
- The last six months in LLMs, illustrated by pelicans on bicyclessimonwillison.net
- GitHub - open-compass/opencompass: OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets. · GitHubgithub.com
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- Paper AI Tigersgleech.org
- State of AI 2025: 100T Token LLM Usage Study | OpenRouteropenrouter.ai
- Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- Building an LLM evaluation framework: best practices | Datadogdatadoghq.com
- LLM-as-a-judge: a complete guide to using LLMs for evaluationsevidentlyai.com