✳flâneur — a map of the web's best reading
Optimizing Mixture of Block Attention
arxiv.org · saved by 1 readers
N/A
Explore this link on the map →related reading
- Sparser Block-Sparse Attention via Token Permutationarxiv.org
- Attention-Residuals/Attention_Residuals.pdf at master · MoonshotAI/Attention-Residuals · GitHubgithub.com
- Future leakage in block-quantized attention | MatXmatx.com
- 2502.11089arxiv.org
- Jacobian Lens – Qwen3.6-27B | Neuronpedianeuronpedia.org
- Overleaf Examplearxiv.org
- Prompt Cache: Modular Attention Reuse for Low-Latency Inferencearxiv.org
- Stand-Alone Self-Attention in Vision Models - NeurIPS-2019-stand-alone-self-attention-in-vision-models-Paper.pdfpapers.nips.cc
- Set Transformer: A Framework for Attention-basedPermutation-Invariant Neural Networksarxiv.org
- GitHub - Masao-Taketani/FOTS_OCR: TensorFlow Implementation of FOTS, Fast Oriented Text Spotting with a Unified Network. · GitHubgithub.com
- GitHub - NVlabs/AutoGaze: AutoGaze automatically removes redundant patches in a video, reducing #tokens in ViT/MLLM by 4x-100x. · GitHubgithub.com
- A User's Guide to FlexAttention in FlashAttention CuTe DSL - Colfax Researchresearch.colfax-intl.com