2403.14123
arxiv.org · 3,999 words · saved by 1 readers
N/A
NOTICE: This arXiv version is an extended version of our paper that has been published in IEEE Micro Journal. AI and Memory Wall Amir Gholami1,2 Zhewei Yao1 Sehoon Kim1 Coleman Hooper1 Michael W. Mahoney1,2,3 Kurt Keutzer1 1 2 3…
related reading
- How To Scale Your Modeljax-ml.github.io
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Making Deep Learning go Brrrr From First Principleshorace.io
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- All About Rooflines | How To Scale Your Modeljax-ml.github.io
- Memory bandwidth constraints imply economies of scale in AI inference — LessWronglesswrong.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- How is LLaMa.cpp possible?finbarr.ca
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- Data movement bottlenecks to large-scale model training: Scaling past 1e28 FLOP | Epoch AIepoch.ai
- AI's Hardware Problemasianometry.substack.com
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com