Omniscaling to MNIST — LessWrong
In this post, I describe a mindset that is flawed, and yet helpful for choosing impactful technical AI safety research projects. The mindset is this: future AI might look very different than AI today, but good ideas are universal. If you want to develop a method that will scale up to powerful future AI systems, your method should also scale down to MNIST. In other words, good ideas omniscale: they work well across all model sizes, domains, and training regimes. Putting the omniscaling mindset into practice is straightforward. Any time you come across a clever-sounding machine learning idea, ask: "can I apply this to MNIST?" If not, then it's not a good idea. If so, run an experiment to see if it works. If it doesn't, then it's not a good idea. If it does, then it might be a good idea, and you can continue as usual to more realistic experiments or theory. In this post, I will: The strongest argument for testing your ideas against MNIST is empirical. Experiments on MNIST have helped to e
x Omniscaling to MNIST — LessWrong Empiricism AI Frontpage 2025 Top Fifty: 12 % 104 Omniscaling to MNIST by cloud 8th Nov 2025 12 min read 3 104 In this post, I describe a mindset that is flawed, and yet helpful for choosing impactful technical AI safety research projects. The mindset is this: future AI might look very different than AI today, but good ideas are universal. If you want to develop a method that will scale up to powerful future AI systems, your method should also scale down to MNIST . In other words, good ideas omniscale : they work well across all model sizes, domains, and train
Explore this link on the map →related reading
- AI 2027ai-2027.com
- The Scaling Hypothesis · Gwern.netgwern.net
- Just Ask for Generalization | Eric Jangevjang.com
- Automated Weak-to-Strong Researcheralignment.anthropic.com
- AI in 2025: gestalt — LessWronglesswrong.com
- I am worried about near-term non-LLM AI developments — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- On neural scaling and the quanta hypothesisericjmichaud.com
- AI 2027ai-2027.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com