✳flâneur — a map of the web's best reading
Full Stack Optimization of Transformer Inference: a Survey
arxiv.org · saved by 1 readers
N/A
Explore this link on the map →related reading
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Efficiently Scaling Transformer Inferencearxiv.org
- The Annotated Transformernlp.seas.harvard.edu
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4. · GitHubgithub.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- GitHub - google-research/tuning_playbook: A playbook for systematically maximizing the performance of deep learning models. · GitHubgithub.com
- DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Modelsarxiv.org
- Neuronpedianeuronpedia.org
- Set Transformer: A Framework for Attention-basedPermutation-Invariant Neural Networksarxiv.org
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attentionarxiv.org
- Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Modelsarxiv.org