flâneur — a map of the web's best reading

Bug Fixes in LLM Training - Gradient Accumulation

unsloth.ai · 2,406 words · saved by 1 readers

Unsloth's Gradient Accumulation fix solves critical errors in LLM Training.

Bug Fixes in LLM Training - Gradient Accumulation unsloth Models Blog Unsloth Studio✨ Docs Blog Bugs in LLM Training - Gradient Accumulation Fix Oct 15, 2024 • By Daniel & Michael Oct 15, 2024 • By Daniel & Michael This past week, we've been fixing a universal issue in Gradient Accumulation that negatively impacts everyone's training, pre-training & finetuning runs for sequence models like LLMs. Unsloth's Gradient Accumulation fix ensures training runs and loss calculations are performed accurately and correctly. The goal of gradient accumulation is to mimic full batch training with reduced VR

Explore this link on the map →

related reading