[2403.06664] Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2403.06664] Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System --> Computer Science > Hardware Architecture arXiv:2403.06664 (cs) [Submitted on 11 Mar 2024] Title: Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System Authors: Hongsun Jang , Jaeyong Song , Jaewon Jung , Jaeyoung Park , Youngsok Kim , Jinho Lee View a PDF of the paper titled Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System, by Hongsun Jang and 5 other authors View PDF HTML (experimental) A
Explore this link on the map →saved by
related reading
- [2312.15159] Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inferencearxiv.org
- [2304.07493] OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantizationarxiv.org
- [2403.00579] NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencingarxiv.org
- How To Scale Your Modeljax-ml.github.io
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- How is LLaMa.cpp possible?finbarr.ca
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B · Hazy Researchhazyresearch.stanford.edu
- Optimizing inference · Hugging Facehuggingface.co
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai