flâneur — a map of the web's best reading

The Unreasonable Effectiveness of Easy Training Data for Hard Tasks

alexandrabarr.beehiiv.com · 1,752 words · saved by 1 readers

Week 2 of Superslow and so far it’s living up to its namesake 🙂 This week, we’re exploring a paper that examines how effective “easy” training data can be in developing strong model capabilities compared to training with more challenging or “hard” data. This paper is authored by one of my best friends, Peter Hase, who recently wrapped up a residency on the safety team at Anthropic. Hope you enjoy! Let’s dive in! Thanks for reading! Hit subscribe to stay updated on the most interesting news in AI. At a high level, this paper aims to answer the question: if I train my models on “easy” data, can I get the same performance as training on “hard” data? Hase et. al measure four things in this paper: How do we know if data is “hard” vs “easy”? Do “easy”-trained models score highly on “hard” tests? What are the tradeoffs between collecting easy vs hard datasets? If the final upshot is that “easy”-trained models perform similarly well to “hard-trained models on hard tests, is this behavior cons

The Unreasonable Effectiveness of Easy Training Data for Hard Tasks 0 Superslow AI Newsletter Posts The Unreasonable Effectiveness of Easy Training Data for Hard Tasks The Unreasonable Effectiveness of Easy Training Data for Hard Tasks Alexandra Barr June 16, 2025 Hi everyone 👋 Week 2 of Superslow and so far it’s living up to its namesake 🙂 This week, we’re exploring a paper that examines how effective “easy” training data can be in developing strong model capabilities compared to training with more challenging or “hard” data. This paper is authored by one of my best friends, Peter Hase, who

Explore this link on the map →

saved by

related reading