The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
Week 2 of Superslow and so far it’s living up to its namesake 🙂 This week, we’re exploring a paper that examines how effective “easy” training data can be in developing strong model capabilities compared to training with more challenging or “hard” data. This paper is authored by one of my best friends, Peter Hase, who recently wrapped up a residency on the safety team at Anthropic. Hope you enjoy! Let’s dive in! Thanks for reading! Hit subscribe to stay updated on the most interesting news in AI. At a high level, this paper aims to answer the question: if I train my models on “easy” data, can I get the same performance as training on “hard” data? Hase et. al measure four things in this paper: How do we know if data is “hard” vs “easy”? Do “easy”-trained models score highly on “hard” tests? What are the tradeoffs between collecting easy vs hard datasets? If the final upshot is that “easy”-trained models perform similarly well to “hard-trained models on hard tests, is this behavior cons
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks 0 Superslow AI Newsletter Posts The Unreasonable Effectiveness of Easy Training Data for Hard Tasks The Unreasonable Effectiveness of Easy Training Data for Hard Tasks Alexandra Barr June 16, 2025 Hi everyone 👋 Week 2 of Superslow and so far it’s living up to its namesake 🙂 This week, we’re exploring a paper that examines how effective “easy” training data can be in developing strong model capabilities compared to training with more challenging or “hard” data. This paper is authored by one of my best friends, Peter Hase, who
saved by
related reading
- The Scaling Hypothesis · Gwern.netgwern.net
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Sporks of AGIsergeylevine.substack.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Ilya Sutskever — We're moving from the age of scaling to the age of researchdwarkesh.com
- Scaling is subtler than it seemsberen.io
- 3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitationalignment.anthropic.com
- Foundation Models for Oversight | Transluce AItransluce.org
- [2602.05910] Chunky Post-Training: Data Driven Failures of Generalizationarxiv.org
- How to Explore to Scale RL Training of LLMs on Hard Problems? – Machine Learning Blog | ML@CMU | Carnegie Mellon Universityblog.ml.cmu.edu
- Why you need to improve your training data, and how to do it << Pete Warden's blogpetewarden.com
- [2509.14786] Pre-training under infinite computearxiv.org