LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by selecting from this list of supported packages. Large language models (LLMs) have been applied in various applications due to their astonishing capabilities. With advancements in technologies such as chain-of-thought (CoT) prompting and in-context learning (ICL), the prompts fed to LLMs are becoming increasingly lengthy, even exceeding tens of thousands of tokens. To accelerate model inference and reduce co
LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, Lili Qiu Microsoft Corporation {hjiang, qianhuiwu, cyl, yuqing.yang, liliqiu}@microsoft.com Abstract Large language models (LLMs) have been applied in various applications due to their astonishing capabilities. With advancements in technologies such as chain-of-thought (CoT) prompting and in-context learning (ICL), the prompts fed to LLMs are becoming increasingly lengthy, even exceeding tens of thousands of tokens. To accelerate model inference and reduce co
Explore this link on the map →saved by
related reading
- LLMs can invent their own compression - Rajan Agarwalrajan.sh
- LongLLMLingua Prompt Compression Guide | LlamaIndexblog.llamaindex.ai
- Composer2.pdfcursor.com
- [2309.10668] Language Modeling Is Compressionarxiv.org
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Can gzip be a language model?nathan.rs
- Optimizing inference · Hugging Facehuggingface.co
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- LLM Resourcesforrestbicker.com
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com