Context Rot: How Increasing Input Tokens Impacts LLM Performance | Chroma Research
Recent developments in LLMs show a trend toward longer context windows, with the input token count of the latest models reaching the millions. Because these models achieve near-perfect scores on widely adopted benchmarks like Needle in a Haystack (NIAH) [1], it’s often assumed that their performance is uniform across long-context tasks. However, NIAH is fundamentally a simple retrieval task, in which a known sentence (the “needle”) is placed in a long document of unrelated text (the “haystack”), and the model is prompted to retrieve it. While scalable, this benchmark typically assesses direct lexical matching, which may not be representative of flexible, semantically oriented tasks. Example Needle in a Haystack (NIAH) Setup We extend the standard NIAH task, to investigate model behavior in previously underexplored settings. We examine the effects of needles with semantic, rather than direct lexical matches, as well as the effects of introducing variations to the haystack content. Addit
Recent developments in LLMs show a trend toward longer context windows, with the input token count of the latest models reaching the millions. Because these models achieve near-perfect scores on widely adopted benchmarks like Needle in a Haystack (NIAH) [ 1 ], it’s often assumed that their performance is uniform across long-context tasks. However, NIAH is fundamentally a simple retrieval task, in which a known sentence (the “needle”) is placed in a long document of unrelated text (the “haystack”), and the model is prompted to retrieve it. While scalable, this benchmark typically assesses direc
related reading
- Recursive Language Models | Alex L. Zhangalexzhang13.github.io
- Extending Context is Hard | kaiokendevkaiokendev.github.io
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- How Long Contexts Faildbreunig.com
- Long Context RAG Performance of LLMs | Databricks Blogdatabricks.com
- GLM-5.2: Built for Long-Horizon Tasksz.ai
- Reimagining LLM Memory: Using Context as Training Data Unlocks Models That Learn at Test-Time | NVIDIA Technical Blogdeveloper.nvidia.com
- gemini_v1_5_report.pdfstorage.googleapis.com
- Recent LLMs can use filler tokens or problem repeats to improve (no-CoT) math performanceblog.redwoodresearch.org
- Generalizing an LLM from 8k to 1M Context using Qwen-Agent | Qwenqwenlm.github.io
- Recursive Language Models: the paradigm of 2026primeintellect.ai