Open Weights isn't Open Training
workshoplabs.ai · 2,924 words · saved by 1 readers
How many monkey-patches does it take to post-train a trillion parameter model?
When I was in college, my data structures professor told a story. It went something like this: "When I was your age, I received an assignment, and encountered an inexplicable bug. I debugged and debugged and found that adding a print statement resolved the bug. I was young like all of you, and I was certain I'd found a bug in the C compiler. Turns out the problem was me." The takeaway was clear: if you have a bug, it's your fault. This is a good heuristic for most cases, but with open source ML infrastructure, you need to throw this advice out the window. There might be features that…
saved by
related reading
- How to Deploy Your Modelhtdym.sailresearch.com
- Composer2.pdfcursor.com
- Google "We Have No Moat, And Neither Does OpenAI"semianalysis.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Keep the Tokens Flowing: Lessons from 16 Open-Source RL Librarieshuggingface.co
- PiTorch: ML on Baremetal Raspberry Pis | projectsmasonjwang.com
- Frontier-scale RL with Kimi K3appliedcompute.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- RL Post-Training on Macs | Pluralis Researchpluralis.ai
- Training great LLMs entirely from ground up in the wilderness as a startup - Yi Tayyitay.net