Test-Time Training with Self-Supervision for Generalization under Distribution Shifts | HTML5
In this paper, we propose Test-Time Training, a general approach for improving the performance of predictive models when training and test data come from different distributions. We turn a single unlabeled test sample into a self-supervised learning problem, on which we update the model parameters before making a prediction. This also extends naturally to data in an online stream. Our simple approach leads to improvements on diverse image classification benchmarks aimed at evaluating robustness to distribution shifts. Supervised learning remains notoriously weak at generalization under distribution shifts. Unless training and test data are drawn from the same distribution, even seemingly minor differences turn out to defeat state-of-the-art models (Recht et al., 2018). Adversarial robustness and domain adaptation are but a few existing paradigms that try to anticipate differences between the training and test distribution with either topological structure or data from the test distribu
Test-Time Training with Self-Supervision for Generalization under Distribution Shifts Yu Sun Xiaolong Wang Zhuang Liu John Miller Alexei A. Efros Moritz Hardt Appendix: Test-Time Training with Self-Supervision for Generalization under Distribution Shifts Yu Sun Xiaolong Wang Zhuang Liu John Miller Alexei A. Efros Moritz Hardt Abstract In this paper, we propose Test-Time Training, a general approach for improving the performance of predictive models when training and test data come from different distributions. We turn a single unlabeled test sample into a self-supervised learning problem, on w
Explore this link on the map →related reading
- Learning with not Enough Data Part 1: Semi-Supervised Learning | Lil'Loglilianweng.github.io
- [2605.01172] A Theory of Generalization in Deep Learningarxiv.org
- arxiv.org/pdf/2511.08544arxiv.org
- [1911.08731] Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalizationarxiv.org
- Just Ask for Generalization | Eric Jangevjang.com
- [1909.02060] Distributionally Robust Language Modelingarxiv.org
- Data Distribution Shifts and Monitoringhuyenchip.com
- Self-Adapting Language Modelsarxiv.org
- Yang Songyang-song.net
- Optimizing LLM Test-Time Compute Involves Solving a Meta-RL Problem – Machine Learning Blog | ML@CMU | Carnegie Mellon Universityblog.ml.cmu.edu
- The Unreasonable Effectiveness of Easy Training Data for Hard Tasksalexandrabarr.beehiiv.com
- Self-Supervised Learning Advances Medical Image Classificationai.googleblog.com