flâneur

Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

arxiv.org · 5,191 words · saved by 1 readers

N/A

Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference Jiaming Tang * 1 2 Yilong Zhao * 1 3 Kan Zhu 3 Guangxuan Xiao 2 Baris Kasikci 3 Song Han 2 4 Abstract As the demand for long-context large language models (LLMs) increases, models with context arXiv:2406.10774v2 [cs.CL] 26 Aug 2024 windows of up to 128K or 1M tokens…

related reading