2306.14048
arxiv.org · 7,834 words · saved by 1 readers
N/A
H2 O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Zhenyu Zhang1 , Ying Sheng2 , Tianyi Zhou3 , Tianlong Chen1 , Lianmin Zheng4 , Ruisi Cai1 , Zhao Song5 , Yuandong Tian6 , Christopher Ré2 , Clark Barrett2 , Zhangyang Wang1 , Beidi Chen6,7 1 University of Texas at Austin, 2 Stanford University, 3…
related reading
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- Cache strategies · Hugging Facehuggingface.co
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- KV Caching Explained: Optimizing Transformer Inference Efficiencyhuggingface.co
- ali (@waterloo_intern) on Xx.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- 2407.15891arxiv.org
- Optimizing inference · Hugging Facehuggingface.co
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inferencearxiv.org