pdoom.org/thesis.html
Neural networks are mean-seeking. They work well when you run inference on data points that lie around the mean of their training data. They embarrassingly fail otherwise. Franz Srambical p(doom) Mihir Mahajan p(doom) Sept. 26, 2024 No DOI yet. Currently, we exploit the paradigm of data-driven self-supervision by using neural networks to approximate the underlying data-generating process of a training distribution, until we run out of data. When we have exhausted the data readily accessible to us, a system trained using naïve self-supervision on such a large-scale dataset will have a reasonably good internal model of neighborhoods in data-space that have a high presence in the training data. The further we go along the long tail of the data distribution, the worse a neural network gets at modeling such data points in its representation-space. Language models are simulators. They will simulate anything, so long as you let the language model ingest enough imitation data at training time.
We pay $300/month to record your screen for AI research. Apply → × --> Neural networks are mean-seeking. They work well when you run inference on data points that lie around the mean of their training data. They embarrassingly fail otherwise. 1 Currently, we exploit the paradigm of data-driven self-supervision by using neural networks to approximate the underlying data-generating process of a training distribution, until we run out of data. When we have exhausted the data readily accessible to us, a system trained using naïve self-supervision on such a large-scale dataset will have a reasonabl
Explore this link on the map →related reading
- A Recipe for Training Neural Networkskarpathy.github.io
- The Only Important Technology Is The Internet - Kevin Lukevinlu.ai
- Bayesian Neural Networkscs.toronto.edu
- A Recipe for Training Neural Networkskarpathy.github.io
- Just Ask for Generalization | Eric Jangevjang.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- The World Inside Neural Networksgoodfire.ai
- The Scaling Hypothesis · Gwern.netgwern.net
- On neural scaling and the quanta hypothesisericjmichaud.com
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Human-like Neural Nets by Catapulting · Gwern.netgwern.net