https://aclanthology.org/P19-1285.pdf
aclanthology.org · 6,350 words · saved by 1 readers
N/A
Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context Zihang Dai⇤12 , Zhilin Yang⇤12 , Yiming Yang1 , Jaime Carbonell1 , Quoc V. Le2 , Ruslan Salakhutdinov1 1 Carnegie Mellon University, 2 Google Brain {dzihang,zhiliny,yiming,jgc,rsalakhu}@cs.cmu.edu, qvl@google.com Abstract Term Memory (LSTM) networks (Hochreiter and…
saved by
related reading
- Seq2seq and Attentionlena-voita.github.io
- transformer_attention.pdfarxiv.org
- Extending Context is Hard | kaiokendevkaiokendev.github.io
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- 1706.03762arxiv.org
- [2512.23675] End-to-End Test-Time Training for Long Contextarxiv.org
- The Transformer Family Version 2.0 | Lil'Loglilianweng.github.io
- Transformers from Scratche2eml.school
- [2011.04006] Long Range Arena: A Benchmark for Efficient Transformersarxiv.org
- [2501.00663] Titans: Learning to Memorize at Test Timearxiv.org
- Transformer (deep learning) - Wikipediaen.wikipedia.org
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com