✳flâneur — a map of the web's best reading
Attention-Residuals/Attention_Residuals.pdf at master · MoonshotAI/Attention-Residuals
github.com · 121 words · saved by 3 readers
N/A
Explore this link on the map →saved by
related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformer (deep learning) - Wikipediaen.wikipedia.org
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- Stand-Alone Self-Attention in Vision Models - NeurIPS-2019-stand-alone-self-attention-in-vision-models-Paper.pdfpapers.nips.cc
- Residual neural network - Wikipediaen.wikipedia.org
- Prompt Cache: Modular Attention Reuse for Low-Latency Inferencearxiv.org
- The Annotated Transformernlp.seas.harvard.edu
- Set Transformer: A Framework for Attention-basedPermutation-Invariant Neural Networksarxiv.org
- pytorch-image-models/timm/models/vision_transformer.py at main · huggingface/pytorch-image-models · GitHubgithub.com
- arxiv.org/pdf/2512.24880#page=3.56arxiv.org
- flash-attention/csrc/flash_attn/src/flash_fwd_kernel.h at main · Dao-AILab/flash-attention · GitHubgithub.com