flâneur — a map of the web's best reading

pdoom.org/thesis.html

pdoom.org · 484 words · saved by 1 readers

Neural networks are mean-seeking. They work well when you run inference on data points that lie around the mean of their training data. They embarrassingly fail otherwise. Franz Srambical p(doom) Mihir Mahajan p(doom) Sept. 26, 2024 No DOI yet. Currently, we exploit the paradigm of data-driven self-supervision by using neural networks to approximate the underlying data-generating process of a training distribution, until we run out of data. When we have exhausted the data readily accessible to us, a system trained using naïve self-supervision on such a large-scale dataset will have a reasonably good internal model of neighborhoods in data-space that have a high presence in the training data. The further we go along the long tail of the data distribution, the worse a neural network gets at modeling such data points in its representation-space. Language models are simulators. They will simulate anything, so long as you let the language model ingest enough imitation data at training time.

We pay $300/month to record your screen for AI research. Apply → × --> Neural networks are mean-seeking. They work well when you run inference on data points that lie around the mean of their training data. They embarrassingly fail otherwise. 1 Currently, we exploit the paradigm of data-driven self-supervision by using neural networks to approximate the underlying data-generating process of a training distribution, until we run out of data. When we have exhausted the data readily accessible to us, a system trained using naïve self-supervision on such a large-scale dataset will have a reasonabl

Explore this link on the map →

related reading