✳flâneur — a map of the web's best reading
arxiv.org/pdf/2512.24880#page=3.56
arxiv.org · 8,721 words · saved by 2 readers
N/A
# link_1eeyhja26ee.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Zhenda Xie; Yixuan Wei; Huanqi Cao; Chenggang Zhao; Chengqi Deng; Jiashi Li; Damai Dai; Huazuo Gao; Jiang Chang; Kuai Yu; Liang Zhao; Shangyan Zhou; Zhean Xu; Zhengyan Zhang; Wangding Zeng; Shengding Hu; Yuqing Wang; Jingyang Yuan; Lean Wang; Wenfeng Liang - Creator=arXiv GenPDF (tex2pdf:57610bf) - Custom.DOI=https://doi.org/10.48550/arXiv.2512.24880 - Custom.License=http://arxiv.org/licenses/nonexclusive-
Explore this link on the map →saved by
related reading
- [2602.05970] Inverse Depth Scaling From Most Layers Being Similararxiv.org
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- How To Scale Your Modeljax-ml.github.io
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- The Practitioner’s Guide to the Maximal Update Parameterization - Cerebrascerebras.ai
- The Practitioner's Guide to the Maximal Update Parameterization | EleutherAI Blogblog.eleuther.ai
- Mamba: The Easy Wayjackcook.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- The Little Book of Deep Learningfleuret.org
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com