[2403.06664] Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2403.06664] Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System --> Computer Science > Hardware Architecture arXiv:2403.06664 (cs) [Submitted on 11 Mar 2024] Title: Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System Authors: Hongsun Jang , Jaeyong Song , Jaewon Jung , Jaeyoung Park , Youngsok Kim , Jinho Lee View a PDF of the paper titled Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System, by Hongsun Jang and 5 other authors View PDF HTML (experimental) A
saved by
related reading
- AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technologyarxiv.org
- [2312.15159] Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inferencearxiv.org
- [2304.07493] OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantizationarxiv.org
- How To Scale Your Modeljax-ml.github.io
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- How is LLaMa.cpp possible?finbarr.ca
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev
- Scalable MatMul-free Language Modelingarxiv.org
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- 1910.02054v3arxiv.org
- Optimizing inference · Hugging Facehuggingface.co