Fetching the paper…
Reading the bibliography…
A major limitation for the broader scope of problems solvable by transformers is the quadratic scaling of computational complexity with input size.
A logical calculus of the ideas immanent in nervous activity
McCulloch, W. S.; and Pitts, W. 1943 · 1943
Earlier work this paper cites.
Kleene. Representation of events in nerve nets and finite automata
Stephen, C. 1956 · 1956
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Werbos, P. J. 1990 · 1990
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Addressing some limitations of transformers with feedback memory
Fan, A.; Lavril, T.; Grave, E.; Joulin, A.; and Sukhbaatar, S. 2020 · 2002
Earlier work this paper cites.
ETC: Encoding Long and Structured Data in Transformers
Ainslie, J.; Ontanon, S.; Alberti, C.; Pham, P.; Ravula, A.; and Sanghai, S. 2020 · 2004
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I.; Peters, M. E.; and Cohan, A. 2020 · 2004
Earlier work this paper cites.
MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph Captioning
Lei, J.; Wang, L.; Shen, Y.; Yu, D.; Berg, T. L.; and Bansal, M. 2020 · 2005
Earlier work this paper cites.
Burtsev, M. S.; Kuratov, Y.; Peganov, A.; and Sapunov, G. V. 2020 · 2006
Earlier work this paper cites.
GMAT: Global memory augmentation for transformers
Gupta, A.; and Berant, J. 2020 · 2006
Earlier work this paper cites.
On the Properties of Neural Machine Translation: Encoder–Decoder Approaches
Cho, K.; van Merriënboer, B.; Bahdanau, D.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Graves, A.; Wayne, G.; and Danihelka, I. 2014 · 2014
Earlier work this paper cites.
The Lean theorem prover (system description)
de Moura, L.; Kong, S.; Avigad, J.; Van Doorn, F.; and von Raumer, J. 2015 · 2015
Earlier work this paper cites.
Learning to Transduce with Unbounded Memory
Grefenstette, E.; Hermann, K. M.; Suleyman, M.; and Blunsom, P. 2015 · 2015
Earlier work this paper cites.
Inferring Algorithmic Patterns with Stack-Augmented Recurrent Nets
Joulin, A.; and Mikolov, T. 2015 · 2015
Earlier work this paper cites.
Sukhbaatar, S.; Szlam, A.; Weston, J.; and Fergus, R. 2015 · 2015
Earlier work this paper cites.
Memory Networks
Weston, J.; Chopra, S.; and Bordes, A. 2015 · 2015
Earlier work this paper cites.
Hybrid computing using a neural network with dynamic external memory
Graves, A.; Wayne, G.; Reynolds, M.; Harley, T.; Danihelka, I.; Grabska-Barwińska, A.; Colmenarejo, S. G.; Grefenstette, E.; Ramalho, T.; Agapiou, J.; Badia, A. P.; Hermann, K. M.; Zwols, Y.; Ostrovski, G.; Cain, A.; King, H.; Summerfield, C.; Blunsom, P.; Kavukcuoglu, K.; and Hassabis, D. 2016 · 2016
Earlier work this paper cites.
Dynamic neural turing machine with soft and hard addressing schemes
Gulcehre, C.; Chandar, S.; Cho, K.; and Bengio, Y. 2016 · 2016
Cited alongside, same era.
Scaling Memory-Augmented Neural Networks with Sparse Reads and Writes
Rae, J. W.; Hunt, J. J.; Harley, T.; Danihelka, I.; Senior, A.; Wayne, G.; Graves, A.; and Lillicrap, T. P. 2016 · 2016
Cited alongside, same era.
Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
Weston, J.; Bordes, A.; Chopra, S.; and Mikolov, T. 2016 · 2016
Cited alongside, same era.
Memory augmented neural networks with wormhole connections
Gulcehre, C.; Chandar, S.; and Bengio, Y. 2017 · 2017
Cited alongside, same era.
Context-aware neural model for temporal information extraction
Meng, Y.; and Rumshisky, A. 2018 · 2018
Recurrent Memory Transformer
Bulatov, A.; Kuratov, Y.; and Burtsev, M. 2022 · 2022
Later among the works it cites.
LongT5: Efficient Text-To-Text Transformer for Long Sequences
Guo, M.; Ainslie, J.; Uthus, D.; Ontanon, S.; Ni, J.; Sung, Y.-H.; and Yang, Y. 2022 · 2022
Later among the works it cites.
Towards a Unified View of Parameter-Efficient Transfer Learning
He, J.; Zhou, C.; Ma, X.; Berg-Kirkpatrick, T.; and Neubig, G. 2022 · 2022
Later among the works it cites.
An empirical analysis of compute-optimal large language model training
Hoffmann, J.; Borgeaud, S.; Mensch, A.; Buchatskaya, E.; Cai, T.; Rutherford, E.; de Las Casas, D.; Hendricks, L. A.; Welbl, J.; Clark, A.; Hennigan, T.; Noland, E.; Millican, K.; van den Driessche, G.; Damoc, B.; Guy, A.; Osindero, S.; Simonyan, K.; Elsen, E.; Vinyals, O.; Rae, J.; and Sifre, L. 2022 · 2022
Later among the works it cites.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; yelong shen; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Transformer-XL: Attentive Language Models beyond a Fixed-Length Context
Dai, Z.; Yang, Z.; Yang, Y.; Carbonell, J.; Le, Q.; and Salakhutdinov, R. 2019 · 2019
Cited alongside, same era.
BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
Star-Transformer
Guo, Q.; Qiu, X.; Liu, P.; Shao, Y.; Xue, X.; and Zhang, Z. 2019 · 2019
Cited alongside, same era.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Cited alongside, same era.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Gao, L.; Biderman, S.; Black, S.; Golding, L.; Hoppe, T.; Foster, C.; Phang, J.; He, H.; Thite, A.; Nabeshima, N.; Presser, S.; and Leahy, C. 2020 · 2020
Cited alongside, same era.
The lean mathematical library
mathlib Community, T. 2020 · 2020
Cited alongside, same era.
Hutchins, D.; Schlag, I.; Wu, Y.; Dyer, E.; and Neyshabur, B. 2022 · 2022
Later among the works it cites.
QuALITY: Question Answering with Long Input Texts, Yes!
Pang, R. Y.; Parrish, A.; Joshi, N.; Nangia, N.; Phang, J.; Chen, A.; Padmakumar, V.; Ma, J.; Thompson, J.; He, H.; and Bowman, S. 2022 · 2022
Later among the works it cites.
Memformer: A Memory-Augmented Transformer for Sequence Modeling
Wu, Q.; Lan, Z.; Qian, K.; Gu, J.; Geramifard, A.; and Yu, Z. 2022a · 2022
Later among the works it cites.
OPT: Open Pre-trained Transformer Language Models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; Mihaylov, T.; Ott, M.; Shleifer, S.; Shuster, K.; Simig, D.; Koura, P. S.; Sridhar, A.; Wang, T.; and Zettlemoyer, L. 2022 · 2022
Later among the works it cites.
CoLT5: Faster Long-Range Transformers with Conditional Computation
Ainslie, J.; Lei, T.; de Jong, M.; Ontañón, S.; Brahma, S.; Zemlyanskiy, Y.; Uthus, D.; Guo, M.; Lee-Thorp, J.; Tay, Y.; Sung, Y.-H.; and Sanghai, S. 2023 · 2023
Closest in time.
Unlimiformer: Long-Range Transformers with Unlimited Length Input
Bertsch, A.; Alon, U.; Neubig, G.; and Gormley, M. R. 2023 · 2023
Closest in time.
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Biderman, S.; Schoelkopf, H.; Anthony, Q.; Bradley, H.; O’Brien, K.; Hallahan, E.; Khan, M. A.; Purohit, S.; Prashanth, U. S.; Raff, E.; Skowron, A.; Sutawika, L.; and van der Wal, O. 2023 · 2023
Closest in time.
Longnet: Scaling transformers to 1,000,000,000 tokens
Ding, J.; Ma, S.; Dong, L.; Zhang, X.; Huang, S.; Wang, W.; and Wei, F. 2023 · 2023
Closest in time.
Lost in the middle: How language models use long contexts
Liu, N. F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; and Liang, P. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
RWKV: Reinventing RNNs for the Transformer Era
Peng, B.; Alcaide, E.; Anthony, Q.; Albalak, A.; Arcadinho, S.; Cao, H.; Cheng, X.; Chung, M.; Grella, M.; GV, K. K.; et al. 2023 · 2023
Closest in time.
Hyena Hierarchy: Towards Larger Convolutional Language Models
Poli, M.; Massaroli, S.; Nguyen, E.; Fu, D. Y.; Dao, T.; Baccus, S.; Bengio, Y.; Ermon, S.; and Re, C. 2023 · 2023
Closest in time.
Retentive network: A successor to transformer for large language models
Sun, Y.; Dong, L.; Huang, S.; Ma, S.; Xia, Y.; Xue, J.; Wang, J.; and Wei, F. 2023 · 2023
Closest in time.