[2410.05686] Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
Abstract:General Purpose Graphics Processing Unit (GPGPU) computing plays a transformative role in deep learning and machine learning by leveraging the computational advantages of parallel processing. Through the power of Compute Unified Device Architecture (CUDA), GPUs enable the efficient execution of complex tasks via massive parallelism. This work explores CPU and GPU architectures, data flow in deep learning, and advanced GPU features, including streams, concurrency, and dynamic parallelism. The applications of GPGPU span scientific computing, machine learning acceleration, real-time rendering, and cryptocurrency mining. This study emphasizes the importance of selecting appropriate parallel architectures, such as GPUs, FPGAs, TPUs, and ASICs, tailored to specific computational tasks and optimizing algorithms for these platforms. Practical examples using popular frameworks such as PyTorch, TensorFlow, and XGBoost demonstrate how to maximize GPU efficiency for training and inference tasks. This resource serves as a comprehensive guide for both beginners and experienced practitioners, offering insights into GPU-based parallel computing and its critical role in advancing machine learning and artificial intelligence.
[2410.05686] Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing --> Computer Science > Distributed, Parallel, and Cluster Computing arXiv:2410.05686 (cs) [Submitted on 8 Oct 2024 ( v1 ), last revised 19 Nov 2025 (this version, v3)] Title: Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing Authors: Ming Li , Ziqian Bi , Tianyang Wang , Yizhu Wen , Qian Niu , Xinyuan Song , Zekun Jiang , Junyu Liu , Benji Peng , Sen Zhang , Xuanhe Pan , Jiawei Xu , Jinlang Wang , Keyu Chen , Caitlyn Heqi Yin , Pohsun Fe
Explore this link on the map →saved by
related reading
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- General-purpose computing on graphics processing units - Wikipediaen.wikipedia.org
- How To Scale Your Modeljax-ml.github.io
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com
- Making Deep Learning go Brrrr From First Principleshorace.io
- The Little Book of Deep Learningfleuret.org
- Pipeline-Parallelism: Distributed Training via Model Partitioningsiboehm.com
- Towards Data Sciencetowardsdatascience.com
- A Hitchhiker’s Guide to ML Training Infrastructure | CMU Software Engineering Institutesei.cmu.edu
- Paradigms of Parallelism | Colossal-AIcolossalai.org
- GPU Performance Background User's Guide - NVIDIA Docsdocs.nvidia.com
- Machine Learning System Resources | std::bodun::blogbodunhu.com