7.pdf
web.stanford.edu · 6,813 words · saved by 1 readers
N/A
Speech and Language Processing. Daniel Jurafsky & James H. Martin. Copyright © 2026. All rights reserved. Draft of August 19, 2026. CHAPTER Transformers and Pretraining 7 “The true art of memory is the art of attention ” Samuel Johnson, Idler #74, September 1759 In this chapter we introduce the transformer, the standard architecture for build- ing large language models, and how to pretrain them and how to use them to generate…
related reading
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformer (deep learning) - Wikipediaen.wikipedia.org
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Everything About Transformerskrupadave.com
- Transformers from Scratche2eml.school
- 9 Transformers – 6.390 - Intro to Machine Learningintroml.mit.edu
- Transformer (deep learning) - Wikipediaen.wikipedia.org
- A Conceptual Guide to Transformers: Part Ibenlevinstein.substack.com
- transformer_attention.pdfarxiv.org
- Transformers from scratch | peterbloem.nlpeterbloem.nl
- The Annotated Transformernlp.seas.harvard.edu