Fetching the paper…
Reading the bibliography…
Unlike recurrent models, conventional wisdom has it that Transformers cannot perfectly model regular languages.
Working memory
Alan D Baddeley and Graham Hitch. 1974 · 1974
Earlier work this paper cites.
Parallel prefix computation
Richard E Ladner and Michael J Fischer. 1980 · 1980
Earlier work this paper cites.
Prefix sums and their applications
Guy E Blelloch. 1990 · 1990
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman. 1990 · 1990
Earlier work this paper cites.
Working memory
Alan Baddeley. 1992 · 1992
Earlier work this paper cites.
Parallel computing using the prefix problem
Sivaramakrishnan Lakshmivarahan and Sudarshan K Dhall. 1994 · 1994
Earlier work this paper cites.
Long-term working memory
K Anders Ericsson and Walter Kintsch. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Serial order: A parallel distributed processing approach
Michael I Jordan. 1997 · 1997
Earlier work this paper cites.
Attention and memory: An integrated framework
Nelson Cowan. 1998 · 1998
Earlier work this paper cites.
Models of working memory
Akira Miyake, Priti Shah, et al. 1999 · 1999
Earlier work this paper cites.
Access to information in working memory: exploring the focus of attention
Klaus Oberauer. 2002 · 2002
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Executive functions
Adele Diamond. 2013 · 2013
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka. 2014 · 2014
Cited alongside, same era.
Wavenet: A generative model for raw audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alexander Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016 · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Theories of working memory: Differences in definition, degree of modularity, role of attention, and purpose
Eryn J Adams, Anh T Nguyen, and Nelson Cowan. 2018 · 2018
Cited alongside, same era.
Parallelizing linear recurrent neural nets over sequence length
Eric Martin and Chris Cundy. 2018 · 2018
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Later among the works it cites.
How many layers and why? An analysis of the model depth in transformers
Antoine Simoulin and Benoit Crabbé. 2021 · 2021
Later among the works it cites.
Adaptive Semiparametric Language Models
Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong. 2021 · 2021
Later among the works it cites.
Exploring length generalization in large language models
Cem Anil, Yuhuai Wu, Anders Johan Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Venkatesh Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur. 2022 · 2022
Later among the works it cites.
Characterizing verbatim short-term memory in neural language models
Kristijan Armeni, Christopher Honey, and Tal Linzen. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Cited alongside, same era.
On the Ability and Limitations of Transformers to Recognize Formal Languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal. 2020 · 2020
Cited alongside, same era.
Depth-adaptive transformer
Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli. 2020 · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Cited alongside, same era.
Overcoming a theoretical limitation of self-attention
David Chiang and Peter Cholak. 2022 · 2022
Later among the works it cites.
The neural data router: Adaptive control flow in transformers improves systematic generalization
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber. 2022 · 2022
Later among the works it cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. 2022 · 2022
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah Smith, and Mike Lewis. 2022 · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022 · 2022
Later among the works it cites.
Neural networks and the chomsky hierarchy
Gregoire Deletang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, and Pedro A Ortega. 2023 · 2023
Closest in time.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. 2023 · 2023
Closest in time.
Resurrecting recurrent neural networks for long sequences
Antonio Orvieto, Samuel L Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De. 2023 · 2023
Closest in time.
Simplified state space layers for sequence modeling
Jimmy T.H. Smith, Andrew Warrington, and Scott Linderman. 2023 · 2023
Closest in time.