[2205.01068] OPT: Open Pre-trained Transformer Language Models
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2205.01068] OPT: Open Pre-trained Transformer Language Models Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computation and Language arXiv:2205.01068 (cs) [Submitted on 2 May 2022 ( v1 ), last revised 21 Jun 2022 (this version, v4)] Title: OPT: Open Pre-trained Transformer Language Models Authors: Susan Zhang , Stephen Roller , Naman Goyal , Mikel Artetxe , Moya Chen , Shuohui Chen , Christopher Dewan , Mona Diab , Xian Li , Xi Victoria Lin , Todor Mihaylov , Myle Ott , Sam Shle
Explore this link on the map →related reading
- Large Language Models Reading List | Sebastian Raschka, PhDsebastianraschka.com
- Transformer Circuits Threadtransformer-circuits.pub
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Understanding Large Language Modelsmagazine.sebastianraschka.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- [2203.15556] Training Compute-Optimal Large Language Modelsarxiv.org
- How do Transformers work? · Hugging Facehuggingface.co
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- Generalized Language Models | Lil'Loglilianweng.github.io
- Recent Advances in Language Model Fine-tuningruder.io
- [2106.09685] LoRA: Low-Rank Adaptation of Large Language Modelsarxiv.org