✳flâneur — a map of the web's best reading
Sparser Block-Sparse Attention via Token Permutation
arxiv.org · saved by 1 readers
N/A
Explore this link on the map →related reading
- 2502.11089arxiv.org
- Optimizing Mixture of Block Attentionarxiv.org
- Attention-Residuals/Attention_Residuals.pdf at master · MoonshotAI/Attention-Residuals · GitHubgithub.com
- Set Transformer: A Framework for Attention-basedPermutation-Invariant Neural Networksarxiv.org
- [2012.09852] SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruningarxiv.org
- Future leakage in block-quantized attention | MatXmatx.com
- Jacobian Lens – Qwen3.6-27B | Neuronpedianeuronpedia.org
- sparse-autoencoders.pdfcdn.openai.com
- GitHub - openai/parameter-golf: Train the smallest LM you can that fits in 16MB. Best model wins! · GitHubgithub.com
- Branches · HazyResearch/intelligence-per-watt · GitHubgithub.com
- Sparse Attention Post-Training for Mechanistic Interpretabilityarxiv.org
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible. · GitHubgithub.com