Prompt Cache: Modular Attention Reuse for Low-Latency Inference
arxiv.org · 7,138 words · saved by 1 readers
N/A
P ROMPT C ACHE : M ODULAR ATTENTION R EUSE FOR L OW-L ATENCY I NFERENCE In Gim 1 Guojun Chen 1 Seung-seob Lee 1 Nikhil Sarda 2 Anurag Khandelwal 1 Lin Zhong 1 A BSTRACT We present Prompt Cache, an approach for accelerating inference for large language models (LLM) by reusing arXiv:2311.04934v2 [cs.CL] 25 Apr 2024…
related reading
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- What is Prompt Caching? | IBMibm.com
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Optimizing inference · Hugging Facehuggingface.co
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- KV Caching Explained: Optimizing Transformer Inference Efficiencyhuggingface.co
- How to make LLMs go fastvgel.me
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Modelsarxiv.org
- Thariq on X: "Lessons from Building Claude Code: Prompt Caching Is Everything " / Xx.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io