Amanda Askell | LinkedIn
I work on aligning and evaluating AI systems, with a focus on large language models. I… · Experience: Anthropic · Education: New York University · Location: San Francisco Bay Area · 418 connections on LinkedIn. View Amanda Askell’s profile on LinkedIn, a professional community of 1 billion members.
Arnoldo Frigessi University of Oslo • 2K followers New paper! Available benchmarks for long-context reasoning in Large Language Models do not distinguish the effect of task complexity, presence of irrelevant information, and task length. @daniel Kaiser invented CogniLoad, a synthetic benchmark grounded in Cognitive Load Theory. CogniLoad is used to test 22 reasoning LLMs, and reveals distinct performance sensitivities, identifying task length as a dominant factor, varied tolerances to intrinsic complexity and a never seen before U-shaped response to presence of distracting information.…
saved by
related reading
- Amanda Askellen.wikipedia.org
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- As Rocks May Think | Eric Jangevjang.com
- Composer2.pdfcursor.com
- Explore | alphaXivalphaxiv.org
- Thoughts on AI in academiatheinfinitesimal.substack.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- machine learning imindslice.substack.com
- 2025: The year in LLMssimonwillison.net
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Group | Sherry Tongshuang Wucs.cmu.edu
- User awareness in frontier modelstransluce.org