flâneur

@ Cutting the Skip: Training Residual-Free Transformers

arxiv.org · 6,217 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Transformers have achieved remarkable success across a wide range of applications, a feat often attributed to their scalability. Yet training them without skip (residual) connections remains notoriously difficult. While skips stabilize optimization, they also disrupt the hierarchical structure of representations, raising the long-standing question of whether transformers can be trained efficiently without them. In this work, we address this problem by analyzing the Jacobian of a skipless transformer block, showing why skips improve conditioning and revealing that their stabilization benefits can be recovered through a principled initialization strategy. Building on this insight, we introduce the first method that enables stable and efficient training of

Yiping JiJames Martens ††thanks: Corresponding e-mail:yiping.ji@adelaide.edu.au Affiliation: Australian Institute for Machine Learning, Adelaide University Affiliation: DATA61, CSIRO Ziqin Zhou Affiliation: Australian Institute for Machine Learning, Adelaide University Peyman Moghadam Affiliation: DATA61, CSIRO Xinyu Zhang Affiliation: University of Auckland Hemanth Saratchandran Affiliation: Australian Institute for Machine Learning, Adelaide University Simon Lucey Affiliation: Australian Institute for Machine Learning, Adelaide University Abstract Transformers have achieved…

saved by

related reading