The Eighty Five Percent Rule for optimal learning | Nature Communications
Is there an optimum difficulty level for training? In this paper, the authors show that for the widely-used class of stochastic gradient-descent based learning algorithms, learning is fastest when the accuracy during training is 85%.
Download PDF Subjects Computer science Human behaviour Learning algorithms Psychology Abstract Researchers and educators have long wrestled with the question of how best to teach their clients be they humans, non-human animals or machines. Here, we examine the role of a single variable, the difficulty of training, on the rate of learning. In many situations we find that there is a sweet spot in which training is neither too easy nor too hard, and where learning progresses most quickly. We derive conditions for this sweet spot for a broad class of learning algorithms in the context of binary cl
related reading
- The 85% Rule for Learning - Scott H Youngscotthyoung.com
- Neural networks and deep learningneuralnetworksanddeeplearning.com
- Discovering 108 tricks to accelerate grokkingkindxiaoming.github.io
- The Little Book of Deep Learningfleuret.org
- The generalization phase diagram — LessWronglesswrong.com
- A critique of pure learning and what artificial neural networks can learn from animal brains | Nature Communicationsnature.com
- A Course in Machine Learningciml.info
- understanding-machine-learning-theory-algorithms.pdfcs.huji.ac.il
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasetsmathai-iclr.github.io
- Learning Beyond Gradientstrinkle23897.github.io
- Human-like Neural Nets by Catapulting · Gwern.netgwern.net
- Investigating the learning coefficient of modular addition: hackathon project — LessWronglesswrong.com