[2001.04063] ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
[2001.04063] ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computation and Language arXiv:2001.04063 (cs) [Submitted on 13 Jan 2020 ( v1 ), last revised 21 Oct 2020 (this version, v3)] Title: ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training Authors: Weizhen Qi , Yu Yan , Yeyun Gong , Dayiheng Liu , Nan Duan , Jiusheng Chen , Ruofei Zhang , Ming Zhou View a PDF of the paper titled
related reading
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- H-Nets - the Past | Goomba Labgoombalab.github.io
- [1409.3215] Sequence to Sequence Learning with Neural Networksarxiv.org
- Introduction to Large Language Models | Machine Learning | Google for Developersdevelopers.google.com
- [2603.05923] Learning Next Action Predictors from Human-Computer Interactionarxiv.org
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- Generalized Language Models | Lil'Loglilianweng.github.io
- Compression and Intelligencegreene.sh
- Can gzip be a language model?nathan.rs
- Dynamic Chunking for End-to-End Hierarchical Sequence Modelingarxiv.org
- Seq2seq and Attentionlena-voita.github.io
- A History of Large Language Modelsgregorygundersen.com