Context Rot: How Increasing Input Tokens Impacts LLM Performance | Chroma Research
Recent developments in LLMs show a trend toward longer context windows, with the input token count of the latest models reaching the millions. Because these models achieve near-perfect scores on widely adopted benchmarks like Needle in a Haystack (NIAH) [1], it’s often assumed that their performance is uniform across long-context tasks. However, NIAH is fundamentally a simple retrieval task, in which a known sentence (the “needle”) is placed in a long document of unrelated text (the “haystack”), and the model is prompted to retrieve it. While scalable, this benchmark typically assesses direct lexical matching, which may not be representative of flexible, semantically oriented tasks. Example Needle in a Haystack (NIAH) Setup We extend the standard NIAH task, to investigate model behavior in previously underexplored settings. We examine the effects of needles with semantic, rather than direct lexical matches, as well as the effects of introducing variations to the haystack content. Addit
Recent developments in LLMs show a trend toward longer context windows, with the input token count of the latest models reaching the millions. Because these models achieve near-perfect scores on widely adopted benchmarks like Needle in a Haystack (NIAH) [ 1 ], it’s often assumed that their performance is uniform across long-context tasks. However, NIAH is fundamentally a simple retrieval task, in which a known sentence (the “needle”) is placed in a long document of unrelated text (the “haystack”), and the model is prompted to retrieve it. While scalable, this benchmark typically assesses direc
Explore this link on the map →related reading
- Recursive Language Models | Alex L. Zhangalexzhang13.github.io
- Extending Context is Hard | kaiokendevkaiokendev.github.io
- Long Context RAG Performance of LLMs | Databricks Blogdatabricks.com
- How Long Contexts Faildbreunig.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Reimagining LLM Memory: Using Context as Training Data Unlocks Models That Learn at Test-Time | NVIDIA Technical Blogdeveloper.nvidia.com
- Recursive Language Models: the paradigm of 2026primeintellect.ai
- gemini_v1_5_report.pdfstorage.googleapis.com
- Generalizing an LLM from 8k to 1M Context using Qwen-Agent | Qwenqwenlm.github.io
- StreamingLLM gives language models unlimited contextbdtechtalks.com
- What is a context window? | IBMibm.com
- Extending the Context of Pretrained LLMs by Dropping their Positional Embeddingspub.sakana.ai