flâneur — a map of the web's best reading

Data Flywheels for LLM Applications

sh-reya.com · 3,985 words · saved by 1 readers

Over the past few months, I have been thinking a lot about workflows to automatically and dynamically improve LLM applications using production data. This stems from our research on validating data quality in LLM pipelines and applications—which is starting to be productionized in both vertical AI applications and LLMOps companies. (I am always very thankful to the teams in industry who find my work useful and are open to collaborating.) My ideas for data flywheels are grounded in several observations: In this post, I’ll outline a three-part framework for approaching this problem. Before we dive into the details of each part, let’s take a look at an overview of the entire LLM application pipeline: Figure 1: Continually-Evolving LLM Application Pipeline, adapted from the source of Figure 1 in this paper. This diagram illustrates my (idealized) architecture of an LLM pipeline, from input processing through evaluation and logging. It showcases ideas I’ll discuss throughout the blog post,

Over the past few months, I have been thinking a lot about workflows to automatically and dynamically improve LLM applications using production data. This stems from our research on validating data quality in LLM pipelines and applications—which is starting to be productionized in both vertical AI applications and LLMOps companies . (I am always very thankful to the teams in industry who find my work useful and are open to collaborating.) My ideas for data flywheels are grounded in several observations: Humans need to be in the loop for evaluation regularly , as human preferences on LLM output

Explore this link on the map →

related reading