2408.14690
arxiv.org · 7,921 words · saved by 1 readers
N/A
Published as a conference paper at ICLR 2025 T RAINING -F REE ACTIVATION S PARSITY IN L ARGE L ANGUAGE M ODELS James Liu1,2∗ Pragaash Ponnusamy2 Tianle Cai3 Han Guo1 Yoon Kim1 Ben Athiwaratkun2 1 2 3 Massachusetts Institute of Technology Together AI Princeton University…
saved by
related reading
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- LLM Resourcesforrestbicker.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- More Efficient In-Context Learning with GLaMblog.research.google
- Efficient LLM inferencefinbarrtimbers.substack.com
- 2502.11089arxiv.org
- Interpretability Infrastructure at Frontier Scale: Harvesting Activations from a Trillion-Parameter Modelgoodfire.ai
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Optimizing inference · Hugging Facehuggingface.co
- Scalable MatMul-free Language Modelingarxiv.org
- Quantization from the ground upngrok.com