Evaluating LLMs is a minefield
normaltech.ai · 214 words · saved by 1 readers
Annotated slides from a recent talk
We have released annotated slides for a talk titled Evaluating LLMs is a minefield. We show that current ways of evaluating chatbots and large language models don't work well, especially for questions about their societal impact. There are no quick fixes, and research is needed to improve evaluation methods. The challenges we highlight are somewhat distinct from those faced by builders of LLMs or by developers interested in comparing between LLMs for adoption. Those challenges are better understood and tackled by evaluation frameworks such as HELM. You can view the annotated slides here.…
related reading
- Evaluating LLMs is a minefieldaisnakeoil.com
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- Successful language model evals - Jason Weijasonwei.net
- gpt-4.pdfcdn.openai.com
- The bitter lesson of LLM evalsparsed.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- LLM evaluation: a beginner's guideevidentlyai.com
- Building an LLM evaluation framework: best practices | Datadogdatadoghq.com
- [2506.13023] A Practical Guide for Evaluating LLMs and LLM-Reliant Systemsarxiv.org
- Things we learned about LLMs in 2024simonwillison.net