flâneur — a map of the web's best reading

9  Transformers – 6.390 - Intro to Machine Learning

introml.mit.edu · 3,420 words · saved by 1 readers

We are actively overhauling the Transformers chapter from the legacy PDF notes to enhance clarity and presentation. Please feel free to raise issues or request more explanation on specific topics. Transformers are a very recent family of architectures that were originally introduced in the field of natural language processing (NLP) in 2017, as an approach to process and understand human language. Since then, they have revolutionized not only NLP but also other domains such as image processing and multi-modal generative AI. Their scalability and parallelizability have made them the backbone of large-scale foundation models, such as GPT, BERT, and Vision Transformers (ViT), powering many state-of-the-art applications. Human language is inherently sequential in nature (e.g., characters form words, words form sentences, and sentences form paragraphs and documents). Prior to the advent of the transformers architecture, recurrent neural networks (RNNs) briefly dominated the field for their a

9 Transformers – 6.390 - Intro to Machine Learning Transformers are a very recent family of architectures that were originally introduced in the field of natural language processing (NLP) in 2017, as an approach to process and understand human language. Since then, they have revolutionized not only NLP but also other domains such as image processing and multi-modal generative AI. Their scalability and parallelizability have made them the backbone of large-scale foundation models, such as GPT, BERT, and Vision Transformers (ViT), powering many state-of-the-art applications. Human language is in

Explore this link on the map →

saved by

related reading