Papers I’ve read this week, Mixture of Experts edition
finbarrtimbers.substack.com · 2,282 words · saved by 2 readers
I read a bunch of papers about conditional routing models
Papers I’ve read this week, Mixture of Experts edition I read a bunch of papers about conditional routing models Finbarr Timbers Aug 04, 2023 40 3 1 Share Papers I’ve read this week, Mixture of Experts edition Mixture of Experts (MoE) models have been getting a lot of attention lately, what with the all the rumours about OpenAI using them in GPT-4. I’ve been reading a lot of the foundational papers about MoE models, and I’ve taken detailed notes, which I wanted to share. This is a bit of a long one, so you might want to read this on the web. Background A standard deep learning model uses the s
saved by
related reading
- Mixture of Experts Explainedhuggingface.co
- [2101.03961] Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsityarxiv.org
- Very Simple MoE Intro1a3orn.com
- [2202.08906] ST-MoE: Designing Stable and Transferable Sparse Expert Modelsar5iv.labs.arxiv.org
- Mixture of experts - Wikipediaen.wikipedia.org
- Mixture of Experts Quantile Balancing: Validated at 32B-A5B (1e22 FLOPs) Scaleopenathena.ai
- [1701.06538] Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layerarxiv.org
- [2209.01667] A Review of Sparse Expert Models in Deep Learningarxiv.org
- Monet: Mixture of Monosemantic Experts for Transformers Explained — LessWronglesswrong.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- DeepSeek-V3: A Large-Scale MoE Pretraining Benchmark for MLPerf Training v6.0mlcommons.org