flâneur — a map of the web's best reading

The huge potential implications of long-context inference

epochai.substack.com · saved by 1 readers

On paper, modern LLMs can ingest many books’ worth of text in one go. For example, Gemini 2.5 Pro has a “context window” of 1 million tokens, enough to stuff in ten copies of Harry Potter and the Philosopher’s Stone.1 But what if we could do lots of inference with much longer contexts? What if LLMs could take in 10 billion tokens of context, and we had the hardware and algorithms to make this usable in practice? The naive use case is being able to take in ever-longer documents. But we think the implications of long context inference could be much greater: It provides an angle of attack on the ability to continually learn new knowledge after the model is deployed, one of the biggest bottlenecks to the real-world utility of current AI systems. It supports a ton of RL scaling: doing more reasoning, verifying model outputs, and generating high-quality RL environments. But there are also bottlenecks. As RL scales to longer runs, research iteration cycles slow down. And you'll also need a lo

On paper, modern LLMs can ingest many books’ worth of text in one go. For example, Gemini 2.5 Pro has a “context window” of 1 million tokens, enough to stuff in ten copies of Harry Potter and the Philosopher’s Stone.1 But what if we could do lots of inference with much longer contexts? What if LLMs could take in 10 billion tokens of context, and we had the hardware and algorithms to make this usable in practice? The naive use case is being able to take in ever-longer documents. But we think the implications of long context inference could be much greater: It provides an angle of attack on the

Explore this link on the map →