✳flâneur — a map of the web's best reading
How Big a Deal are MatMul-Free Transformers? — LessWrong
lesswrong.com · 2,518 words · saved by 1 readers
If you’re already familiar with the technical side of LLMs, you can skip the first section. …
x How Big a Deal are MatMul-Free Transformers? — LessWrong Academic Papers Transformers AI Frontpage 19 How Big a Deal are MatMul-Free Transformers? by JustisMills 27th Jun 2024 Linkpost for justismills.substack.com 6 min read 6 19 If you’re already familiar with the technical side of LLMs, you can skip the first section. The story so far Modern Large Language Models - your ChatGPTs, your Geminis - are a particular kind of transformer , a deep learning architecture invented about seven years ago . Without getting into the weeds , transformers basically work by turning an input into numbers, an
Explore this link on the map →related reading
- How To Scale Your Modeljax-ml.github.io
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Bits, FLOPS, and Watts: A Systems-Level Perspective of Scaling LLMs — Part 1 | by Asheesh Goja | Mediummedium.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Transformers from Scratche2eml.school
- All the Transformer Math You Need to Know | How To Scale Your Modeljax-ml.github.io
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- Linear Transformers Are Faster After All – Manifest AImanifestai.com
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- trees are harlequins, words are harlequins - I don't think you're drawing the right lesson from...nostalgebraist.tumblr.com