[Hero Run] 535B-A23B on 18T tokens · Issue #8435 · marin-community/marin
Issue is to describe the Hero Run scaling ladder, config, and track its performance. Tracking: https://wandb.ai/marin-community/marin_moe/reports/535B-A23B-18T-Token-Hero-Run-Scaling-Ladder--Vmlldz...
Issue is to describe the Hero Run scaling ladder, config, and track its performance. Tracking: https://wandb.ai/marin-community/marin_moe/reports/535B-A23B-18T-Token-Hero-Run-Scaling-Ladder--VmlldzoxNzc2MDM5Ng Details on the full model spec is in the comment below. Why run a scaling ladder? This lets us project the model performance and compare it to the scaling projections of our prior recipes. If our projection looks meaningfully worse, then we know we need to go back to the drawing board on some aspect of the model or data. This step can catch a lot of bugs! We can look at how…
saved by
related reading
- Scaling Laws That Extrapolate 300× Past the Fitopenathena.ai
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- Marin 535B-A23B launch noteopenathena.ai
- Composer2.pdfcursor.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- On neural scaling and the quanta hypothesisericjmichaud.com
- 2404.10102v1arxiv.org
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- 10x Data Efficiency - NanoGPT Slowrunqlabs.sh
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com