deAI - Part 2: Decentralized Training
Special thanks to Sam Kim, the team at FortyTwo Network (Vlad Larin, Alex Firsov, Ivan Nikitin), Travis Good (Ambient), Alexander Long (Pluralis), Dillon Rolnick (Nous Research) for incredibly valuable discussions, insights and feedback of this report. If you are unfamiliar with the premise of how transformers work in the context of Machine Learning, please read through “Transformers 101” in order to better understand the rest of this report. Feel free to skip around the report as you choose, most sections are agnostic to each other until it all comes together at the very end of the report. The multi-headed attention mechanism in transformers, which scales linearly with context length, enables the processing of vast amounts of data in parallel. This groundbreaking approach, while highly efficient, demands extraordinarily large amounts of computational resources. Training large-scale models such as Llama 3 has already pushed boundaries by utilizing thousands of GPUs—estimates suggest a
Explore this link on the map →