[2310.09949] Chameleon: a heterogeneous and disaggregated accelerator system for retrieval-augmented language models
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
View PDF HTML (experimental) Abstract:A Retrieval-Augmented Language Model (RALM) combines a large language model (LLM) with a vector database to retrieve context-specific knowledge during text generation. This strategy facilitates impressive generation quality even with smaller models, thus reducing computational demands by orders of magnitude. To serve RALMs efficiently and flexibly, we propose Chameleon, a heterogeneous accelerator system integrating both LLM and vector search accelerators in a disaggregated architecture. The heterogeneity ensures efficient serving for both inference and…
saved by
related reading
- AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technologyarxiv.org
- [2312.15159] Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inferencearxiv.org
- [2403.00579] NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencingarxiv.org
- Quantization and Hardware Architecture Co-Design for Matrix-Vector Multiplications of Large Language Modelsieeexplore.ieee.org
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Recursive Language Models | Alex L. Zhangalexzhang13.github.io
- Hermes: Algorithm-System Co-design for Efficient Retrieval-Augmented Generation At-Scalemichaeltshen.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- What is Retrieval Augmented Generation (RAG)? | Databricksdatabricks.com
- Retrieval-Augmented Generation for Large Language Models: A Surveyarxiv.org
- Retrieval Augmented Generation: Streamlining the creation of intelligent natural language processing modelsai.facebook.com
- Towards Data Sciencetowardsdatascience.com