Data Flywheels for LLM Applications
Over the past few months, I have been thinking a lot about workflows to automatically and dynamically improve LLM applications using production data. This stems from our research on validating data quality in LLM pipelines and applications—which is starting to be productionized in both vertical AI applications and LLMOps companies. (I am always very thankful to the teams in industry who find my work useful and are open to collaborating.) My ideas for data flywheels are grounded in several observations: In this post, I’ll outline a three-part framework for approaching this problem. Before we dive into the details of each part, let’s take a look at an overview of the entire LLM application pipeline: Figure 1: Continually-Evolving LLM Application Pipeline, adapted from the source of Figure 1 in this paper. This diagram illustrates my (idealized) architecture of an LLM pipeline, from input processing through evaluation and logging. It showcases ideas I’ll discuss throughout the blog post,
Over the past few months, I have been thinking a lot about workflows to automatically and dynamically improve LLM applications using production data. This stems from our research on validating data quality in LLM pipelines and applications—which is starting to be productionized in both vertical AI applications and LLMOps companies . (I am always very thankful to the teams in industry who find my work useful and are open to collaborating.) My ideas for data flywheels are grounded in several observations: Humans need to be in the loop for evaluation regularly , as human preferences on LLM output
Explore this link on the map →related reading
- Building an LLM evaluation framework: best practices | Datadogdatadoghq.com
- LLM evaluation: a beginner's guideevidentlyai.com
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Evaluating LLM Applicationshumanloop.com
- What We Learned from a Year of Building with LLMs (Part II) – O’Reillyoreilly.com
- Your AI Product Needs Evals – Hamel's Blog - Hamel Husainhamel.dev
- Building LLM applications for productionhuyenchip.com
- LLM-as-a-judge: a complete guide to using LLMs for evaluationsevidentlyai.com
- A pragmatic guide to LLM evals for devsnewsletter.pragmaticengineer.com
- The bitter lesson of LLM evalsparsed.com