flâneur — a map of the web's best reading

Everything about Distributed Training and Efficient Finetuning | Sumanth's Personal Website

sumanthrh.com · 8,616 words · saved by 1 readers

A deep dive into distributed training and efficient finetuning - DeepSpeed ZeRO, FSDP, practical guidelines and gotchas with multi-GPU and multi-node training

Everything about Distributed Training and Efficient Finetuning A deep dive into distributed training and efficient finetuning - DeepSpeed ZeRO, FSDP, practical guidelines and gotchas with multi-GPU and multi-node training Last updated on Jan 19, 2024 40 min read There's been an insane amount of interest in large language models (LLMs) these days, with a very special open source community of hackers figuring out the best way to finetune, serve and run inference on consumer-grade hardware. A number of excellent open-source codebases have popped up to meet these needs, notably FastChat , Axolotl

Explore this link on the map →

related reading