flâneur

vatsal0.github.io/blog/emergence.html

vatsal0.github.io · saved by 1 readers

Language model capabilities appear at random, unpredictable moments in training. When they do appear, the improvement is sharp and abrupt. Both behaviors trace back to a single bottleneck: learning which tokens attention should prioritize. June 24, 2026 No DOI yet. Vatsal Baherwani‡ Zixi Chen Shikai Qiu Andrew Gordon Wilson Pavel Izmailov New York University June 25, 2026 ‡ Correspondence to vatsalbaherwani@nyu.edu Paper Neural scaling trends for pretraining performance are surprisingly predictable. But in reality we don't judge LLMs by their pretraining loss; rather, we care about whether they can solve practical downstream tasks. And these practical capabilities, historically, have not been so predictable. For example, GPT-3 resembled a step-function-like jump in in-context learning capability that no one saw coming. Many such emergent abilities have been widely documented in the past few years: as we continue to scale LLMs, new capabilities such as in-context learning, question answ

saved by