Chinchilla
arxiv.org · 6,259 words · saved by 1 readers
N/A
Training Compute-Optimal Large Language Models Jordan Hoffmann★, Sebastian Borgeaud★, Arthur Mensch★, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals and Laurent…
saved by
related reading
- [2203.15556] Training Compute-Optimal Large Language Modelsarxiv.org
- How To Scale Your Modeljax-ml.github.io
- Training a compute-optimal gpt2-small – Tomek Korbak — personal homepagetomekkorbak.com
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- IsoFLOP curves of large language models are extremely flatseverelytheoretical.wordpress.com
- chinchilla's wild implications — LessWronglesswrong.com
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Publications — Google DeepMinddeepmind.com
- chinchilla's wild implications — AI Alignment Forumalignmentforum.org
- 2404.10102v1arxiv.org