2305.13245
arxiv.org · 3,525 words · saved by 1 readers
N/A
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints Joshua Ainslie∗, James Lee-Thorp∗, Michiel de Jong∗ † Yury Zemlyanskiy, Federico Lebrón, Sumit Sanghai Google Research Abstract show…
related reading
- [2305.13245] GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpointsarxiv.org
- 1911.02150arxiv.org
- Multi Query Attention (MQA) and Grouped-Query Attention (GQA)tinkerd.net
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- Mamba: The Easy Wayjackcook.com
- A Gentle Introduction to Multi-Head Latent Attention (MLA) - MachineLearningMastery.commachinelearningmastery.com
- 1706.03762arxiv.org
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- ali (@waterloo_intern) on Xx.com