flâneur — a map of the web's best reading

Ambiguous out-of-distribution generalization on an algorithmic task — LessWrong

lesswrong.com · 4,790 words · saved by 1 readers

Introduction It's now well known that simple neural network models often "grok" algorithmic tasks. That is, when trained for many epochs on a subset…

x Ambiguous out-of-distribution generalization on an algorithmic task — LessWrong Deceptive Alignment Distributional Shifts Grokking (ML) Interpretability (ML & AI) MATS Program AI Frontpage 84 Ambiguous out-of-distribution generalization on an algorithmic task by Wilson Wu , Louis Jaburi 13th Feb 2025 13 min read 6 84 Introduction It's now well known that simple neural network models often "grok" algorithmic tasks. That is, when trained for many epochs on a subset of the full input space, the model quickly attains perfect train accuracy and then, much later, near-perfect test accuracy. In the

Explore this link on the map →

related reading