[2302.05442] Scaling Vision Transformers to 22 Billion Parameters
Abstract:The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Vision Transformers (ViT) have introduced the same architecture to image and video modelling, but these have not yet been successfully scaled to nearly the same degree; the largest dense ViT contains 4B parameters (Chen et al., 2022). We present a recipe for highly efficient and stable training of a 22B-parameter ViT (ViT-22B) and perform a wide variety of experiments on the resulting model. When evaluated on downstream tasks (often with a lightweight linear model on frozen features), ViT-22B demonstrates increasing performance with scale. We further observe other interesting benefits of scale, including an improved tradeoff between fairness and performance, state-of-the-art alignment to human visual perception in terms of shape/texture bias, and improved robustness. ViT-22B demonstrates the potential for "LLM-like" scaling in vision, and provides key steps towards getting there.
[2302.05442] Scaling Vision Transformers to 22 Billion Parameters --> Computer Science > Computer Vision and Pattern Recognition arXiv:2302.05442 (cs) [Submitted on 10 Feb 2023] Title: Scaling Vision Transformers to 22 Billion Parameters Authors: Mostafa Dehghani , Josip Djolonga , Basil Mustafa , Piotr Padlewski , Jonathan Heek , Justin Gilmer , Andreas Steiner , Mathilde Caron , Robert Geirhos , Ibrahim Alabdulmohsin , Rodolphe Jenatton , Lucas Beyer , Michael Tschannen , Anurag Arnab , Xiao Wang , Carlos Riquelme , Matthias Minderer , Joan Puigcerver , Utku Evci , Manoj Kumar , Sjoerd van S
Explore this link on the map →saved by
related reading
- How To Scale Your Modeljax-ml.github.io
- [2605.05331] ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parametersarxiv.org
- The Scaling Hypothesis · Gwern.netgwern.net
- 2403.09611.pdfarxiv.org
- [2106.03348] ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Biasarxiv.org
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- Large Language Models Reading List | Sebastian Raschka, PhDsebastianraschka.com
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- On the speed of ViTs and CNNslucasb.eyer.be
- [1909.08053] Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelismarxiv.org
- Bits, FLOPS, and Watts: A Systems-Level Perspective of Scaling LLMs — Part 1 | by Asheesh Goja | Mediummedium.com
- [2201.03545] A ConvNet for the 2020sarxiv.org