[2401.11459] AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology
Pi Day is Giving Day: arXiv depends on donations to operate and keep science open for all. Give back to arXiv on 3.14.24! Help | Advanced Search arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
View PDF HTML (experimental) Abstract:Large language models (LLMs) with Transformer architectures have become phenomenal in natural language processing, multimodal generative artificial intelligence, and agent-oriented artificial intelligence. The self-attention module is the most dominating sub-structure inside Transformer-based LLMs. Computation using general-purpose graphics processing units (GPUs) inflicts reckless demand for I/O bandwidth for transferring intermediate calculation results between memories and processing units. To tackle this challenge, this work develops a fully…
saved by
related reading
- [2403.00579] NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencingarxiv.org
- [2312.15159] Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inferencearxiv.org
- [2403.06664] Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real Systemarxiv.org
- FTRANS: Energy-Efficient Acceleration of Transformers using FPGAarxiv.org
- Quantization and Hardware Architecture Co-Design for Matrix-Vector Multiplications of Large Language Modelsieeexplore.ieee.org
- Chameleon: a Heterogeneous and Disaggregated Accelerator System for Retrieval-Augmented Language Modelsarxiv.org
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Can LLMs Be Computers?percepta.ai
- MatX: High-throughput chips for LLMsmatx.com
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu