✳flâneur — a map of the web's best reading
Papers I’ve read this week, Mixture of Experts edition
finbarrtimbers.substack.com · 2,282 words · saved by 2 readers
I read a bunch of papers about conditional routing models
Papers I’ve read this week, Mixture of Experts edition I read a bunch of papers about conditional routing models Finbarr Timbers Aug 04, 2023 40 3 1 Share Papers I’ve read this week, Mixture of Experts edition Mixture of Experts (MoE) models have been getting a lot of attention lately, what with the all the rumours about OpenAI using them in GPT-4. I’ve been reading a lot of the foundational papers about MoE models, and I’ve taken detailed notes, which I wanted to share. This is a bit of a long one, so you might want to read this on the web. Background A standard deep learning model uses the s
Explore this link on the map →saved by
related reading
- Very Simple MoE Intro1a3orn.com
- [2202.08906] ST-MoE: Designing Stable and Transferable Sparse Expert Modelsar5iv.labs.arxiv.org
- Mixture of experts - Wikipediaen.wikipedia.org
- Monet: Mixture of Monosemantic Experts for Transformers Explained — LessWronglesswrong.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- [1701.06538] Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layerarxiv.org
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- An Alternative to Test-Time Scalingrentry.org
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- How To Scale Your Modeljax-ml.github.io
- Composer2.pdfcursor.com