What Can Transformers Learn In-Context?
arxiv.org · 7,820 words · saved by 1 readers
N/A
What Can Transformers Learn In-Context? A Case Study of Simple Function Classes Shivam Garg* Dimitris Tsipras* Stanford University Stanford University shivamg@cs.stanford.edu tsipras@stanford.edu arXiv:2208.01066v3 [cs.CL] 11 Aug 2023…
related reading
- [2212.07677] Transformers learn in-context by gradient descentarxiv.org
- In-context Learning and Induction Headstransformer-circuits.pub
- How does in-context learning work? A framework for understanding the differences from traditional supervised learning | SAIL Blogai.stanford.edu
- Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Modelsarxiv.org
- In-context learning creates task vectorsarxiv.org
- Learning without training: The implicit dynamics of in-context learningarxiv.org
- Uncovering mesa-optimization algorithms in Transformersarxiv.org
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?arxiv.org
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- [2402.01258] Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscapearxiv.org
- Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thoughtarxiv.org