New fast transformer inference ASIC — Sohu by Etched — LessWrong
Electrical engineer here. I read the publicity statement, and from my point of view it is both (a) a major advance, if true, and (b) entirely plausible. When you switch from a programmable device (e.g. GPU) to a similarly sized special purpose ASIC, it is not unreasonable to pick up a factor of 10 to 50 in performance. The tradeoff is that the GPU can do many more things than the ASIC, and the ASIC takes years to design. They claim they started design in 2022, on a transformer-only device, on the theory that transformers were going to be popular. And boy, did they luck out. I don‘t know if other people can tell, but to me, that statement oozes with engineering glee. They’re so happy! I would love to see a technical paper on how they did it. Of course they may be lying. One 8xSohu server replaces 160 H100 GPUs. Benchmarks are for Llama-3 70B in FP8 precision, 2048 input/128 output lengths. What would happen if AI models get 20x faster and cheaper overnight? So there is an obliq
x New fast transformer inference ASIC — Sohu by Etched — LessWrong AI Personal Blog 9 New fast transformer inference ASIC — Sohu by Etched by lemonhope 26th Jun 2024 1 min read 9 9 This is a linkpost for https://www.etched.com/announcing-etched I would bet that ASICs will run the roost in a few years and this is only the beginning. They claim 500k tokens per second with Llama 70B. Seems to be exactly what it looks like, an ASIC. Curious if this is somehow not what it looks like. New fast transformer inference ASIC — Sohu by Etched 18 Carl Feynman 6 Adrian Kelly 2 Tao Lin 5 Vladimir_Nesov 1 Has
Explore this link on the map →saved by
related reading
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Stop trying to make Etched happen. - zach's tech blogzach.be
- Reiner Pope of MatX on accelerating AI with transformer-optimized chipscheekypint.substack.com
- How To Scale Your Modeljax-ml.github.io
- An Interview with MatX CEO Reiner Pope About LLM Chipschipstrat.com
- Making Deep Learning go Brrrr From First Principleshorace.io
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- How is LLaMa.cpp possible?finbarr.ca
- The Inference Shift – Stratechery by Ben Thompsonstratechery.com
- Memory bandwidth constraints imply economies of scale in AI inference — LessWronglesswrong.com
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- The Architecture of Dominance: NVIDIA’s Rubin CPX and the $254 Billion Inference Warsshanakaanslemperera.substack.com