✳flâneur — a map of the web's best reading
Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and Parallelization
usenix.org · saved by 1 readers
N/A
Explore this link on the map →related reading
- Deriving Muonjeremybernste.in
- GitHub - google-research/tuning_playbook: A playbook for systematically maximizing the performance of deep learning models. · GitHubgithub.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- The Little Book of Deep Learningfleuret.org
- Annotated Research Paper Implementations: Transformers, StyleGAN, Stable Diffusion, DDPM/DDIM, LayerNorm, Nucleus Sampling and morenn.labml.ai
- CVPR2023_eff_tutorial_molchanov.pdfnvlabs.github.io
- DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Modelsarxiv.org
- GitHub - karpathy/autoresearch: AI agents running research on single-GPU nanochat training automatically · GitHubgithub.com
- 2404.17625arxiv.org
- The Practitioner’s Guide to the Maximal Update Parameterization - Cerebrascerebras.ai
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decodingarxiv.org
- GitHub - labmlai/annotated_deep_learning_paper_implementations: 🧑🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adagithub.com