arxiv.org/pdf/2512.24880#page=3.56
arxiv.org · 8,721 words · saved by 2 readers
N/A
# link_1eeyhja26ee.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Zhenda Xie; Yixuan Wei; Huanqi Cao; Chenggang Zhao; Chengqi Deng; Jiashi Li; Damai Dai; Huazuo Gao; Jiang Chang; Kuai Yu; Liang Zhao; Shangyan Zhou; Zhean Xu; Zhengyan Zhang; Wangding Zeng; Shengding Hu; Yuqing Wang; Jingyang Yuan; Lean Wang; Wenfeng Liang - Creator=arXiv GenPDF (tex2pdf:57610bf) - Custom.DOI=https://doi.org/10.48550/arXiv.2512.24880 - Custom.License=http://arxiv.org/licenses/nonexclusive-
saved by
related reading
- [2512.24880] mHC: Manifold-Constrained Hyper-Connectionsarxiv.org
- mHC: Manifold-Constrained Hyper-Connectionsalphaxiv.org
- [2602.05970] Inverse Depth Scaling From Most Layers Being Similararxiv.org
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Hyper-Connectionsarxiv.org
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Cutting the Skip: Training Residual-Free Transformersarxiv.org
- DeepSeek and the Day Before New Year'swheremachinesthink.substack.com
- The Practitioner’s Guide to the Maximal Update Parameterization - Cerebrascerebras.ai
- [2607.13491] DeepLoop: Depth Scaling for Looped Transformersarxiv.org
- bachlechner21a.pdfproceedings.mlr.press