Evaluating LLM Applications
An ever-increasing number of companies are using large language models (LLMs) to transform both their product experiences and internal operations. These kinds of foundation models represent a new computing platform. The process of prompt engineering is replacing aspects of software development and the scope of what software can achieve is rapidly expanding. In order to effectively leverage LLMs in production, having confidence in how they perform is paramount. This represents a unique challenge for most companies given the inherent novelty and complexities surrounding LLMs. Unlike traditional software and non-generative machine learning (ML) models, evaluation is subjective, hard to automate and the risk of the system going embarrassingly wrong is higher. This post provides some thoughts on evaluating LLMs and discusses some emerging patterns I've seen work well in practice from experience with thousands of teams deploying LLM applications in production. It’s important to first underst
Blog post type Blog Published on 2024/02/06 Evaluating LLM Applications By Peter Hayes Cofounder and CTO By Peter Hayes Cofounder and CTO An ever-increasing number of companies are using large language models (LLMs) to transform both their product experiences and internal operations. These kinds of foundation models represent a new computing platform. The process of prompt engineering is replacing aspects of software development and the scope of what software can achieve is rapidly expanding. In order to effectively leverage LLMs in production, having confidence in how they perform is paramoun
Explore this link on the map →related reading
- LLM evaluation: a beginner's guideevidentlyai.com
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- The bitter lesson of LLM evalsparsed.com
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- Building an LLM evaluation framework: best practices | Datadogdatadoghq.com
- LLM-as-a-judge: a complete guide to using LLMs for evaluationsevidentlyai.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- A pragmatic guide to LLM evals for devsnewsletter.pragmaticengineer.com
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- What We Learned from a Year of Building with LLMs (Part II) – O’Reillyoreilly.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Demystifying evals for AI agents \ Anthropicanthropic.com