Reframing Reinforcement Learning as Sequence Modeling with Transformers?
The Transformer Network, developed by Google and presented in a NeurIPS2017 paper, is one of the few papers that can truly claim to havefundamentally transfo...
The Transformer Network , developed by Google and presented in a NeurIPS 2017 paper , is one of the few papers that can truly claim to have fundamentally transformed (pun intended) the field of Artificial Intelligence. Transformer Networks have become the foundation of some of the most dramatic performance advances in Natural Language Processing (NLP). Two prominent examples are Google’s BERT model , which uses a bidirectional Transformer, and OpenAI’s line of GPT models , which uses a unidirectional Transformer. Both papers have substantially helped out their respective companies’ bottom line
Explore this link on the map →related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformers from Scratche2eml.school
- transformer_attention.pdfarxiv.org
- The Annotated Transformernlp.seas.harvard.edu
- Large Language Models Reading List | Sebastian Raschka, PhDsebastianraschka.com
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- The Topological Trouble With Transformersarxiv.org
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Efficient World Models with Context-Aware Tokenizationarxiv.org
- Transformers from Scratchbrandonrohrer.com
- Transformer (deep learning) - Wikipediaen.wikipedia.org
- LLM Resourcesforrestbicker.com