Successful language model evals — Jason Wei
Everybody uses evaluation benchmarks (“evals”), but I think they deserve more attention than they are currently getting. Evals are incentives for the research community, and breakthroughs are often closely linked to a huge performance jump on some eval. In fact, I’d argue that a key job of the team
Successful language model evals May 24 Written By Jason Wei Everybody uses evaluation benchmarks (“evals”), but I think they deserve more attention than they are currently getting. Evals are incentives for the research community, and breakthroughs are often closely linked to a huge performance jump on some eval. In fact, I’d argue that a key job of the team lead is to dictate what eval to optimize. What is the definition of a successful eval? I’d say that if an eval is used in breakthrough papers and trusted within the community, then it’s clearly successful. Here are some of the successful ev
saved by
related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- The bitter lesson of LLM evalsparsed.com
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- A starter guide for evals — AI Alignment Forumalignmentforum.org
- A statistical approach to model evaluations \ Anthropicanthropic.com
- LLM evaluation: a beginner's guideevidentlyai.com
- [2411.00640] Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluationsarxiv.org
- A starter guide for evals — LessWronglesswrong.com
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org