LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by selecting from this list of supported packages. Large language models (LLMs) have been applied in various applications due to their astonishing capabilities. With advancements in technologies such as chain-of-thought (CoT) prompting and in-context learning (ICL), the prompts fed to LLMs are becoming increasingly lengthy, even exceeding tens of thousands of tokens. To accelerate model inference and reduce co
LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, Lili Qiu Microsoft Corporation {hjiang, qianhuiwu, cyl, yuqing.yang, liliqiu}@microsoft.com Abstract Large language models (LLMs) have been applied in various applications due to their astonishing capabilities. With advancements in technologies such as chain-of-thought (CoT) prompting and in-context learning (ICL), the prompts fed to LLMs are becoming increasingly lengthy, even exceeding tens of thousands of tokens. To accelerate model inference and reduce co
saved by
related reading
- LLMs can invent their own compression - Rajan Agarwalrajan.sh
- Compression and Intelligencegreene.sh
- LongLLMLingua Prompt Compression Guide | LlamaIndexblog.llamaindex.ai
- [2309.10668] Language Modeling Is Compressionarxiv.org
- Composer2.pdfcursor.com
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4.github.com
- Quantization from the ground upngrok.com
- Productizing Large Language Modelsblog.replit.com
- Efficient LLM inferencefinbarrtimbers.substack.com
- Compression is predictionngrok.com
- Can gzip be a language model?nathan.rs
- LLM Optimization via Synthetic Distillationanarchyai.substack.com