Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models
arxiv.org · 5,903 words · saved by 1 readers
N/A
Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models Steve Yadlowsky, Lyric Doshi, Nilesh Tripuraneni {yadlowsky, lyric, nileshtrip}@google.com Google DeepMind arXiv:2311.00871v1 [cs.LG] 1 Nov 2023 November 3,…
related reading
- In-context Learning and Induction Headstransformer-circuits.pub
- What Can Transformers Learn In-Context?arxiv.org
- Transformer Circuits Threadtransformer-circuits.pub
- How does in-context learning work? A framework for understanding the differences from traditional supervised learning | SAIL Blogai.stanford.edu
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- [2212.07677] Transformers learn in-context by gradient descentarxiv.org
- In-context learning creates task vectorsarxiv.org
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- radford2018improving.pdfcs.ubc.ca
- Uncovering mesa-optimization algorithms in Transformersarxiv.org
- Generalized Language Models | Lil'Loglilianweng.github.io