Fetching the paper…
Reading the bibliography…
Transformers have supplanted recurrent models in a large number of NLP tasks.
On the computational power of rnns
Samuel A Korsky and Robert C Berwick. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Counter machines and counter languages
Patrick C Fischer, Albert R Meyer, and Arnold L Rosenberg. 1968 · 1968
Earlier work this paper cites.
Formal language theory: refining the chomsky hierarchy
Gerhard Jäger and James Rogers. 2012 · 1970
Earlier work this paper cites.
Dot-depth of star-free events
Rina Cohen and Janusz Brzozowski. 1971 · 1971
Earlier work this paper cites.
Counter-Free Automata (M.I.T. Research Monograph No. 65)
Robert McNaughton and Seymour A. Papert. 1971 · 1971
Earlier work this paper cites.
Finite Automata, Formal Logic, and Circuit Complexity
Howard Straubing. 1994 · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Lstm recurrent networks learn simple context-free and context-sensitive languages
Felix A Gers and E Schmidhuber. 2001 · 2001
Earlier work this paper cites.
A field guide to dynamical recurrent networks
John F Kolen and Stefan C Kremer. 2001 · 2001
Earlier work this paper cites.
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2002
Earlier work this paper cites.
Learning to encode position for transformer with continuous dynamical model
Xuanqing Liu, Hsiang-Fu Yu, Inderjit Dhillon, and Cho-Jui Hsieh. 2020 · 2003
Earlier work this paper cites.
On the linguistic capacity of real-time counter automata
William Merrill. 2020 · 2004
Earlier work this paper cites.
A formal hierarchy of rnn architectures
William Merrill, Gail Weiss, Yoav Goldberg, Roy Schwartz, Noah A Smith, and Eran Yahav. 2020 · 2004
Cited alongside, same era.
Pretraining on non-linguistic structure as a tool for analyzing learning bias in language models
Isabel Papadimitriou and Dan Jurafsky. 2020 · 2004
Cited alongside, same era.
On the computational power of transformers and its implications in sequence modeling
Satwik Bhattamishra, Arkil Patel, and Navin Goyal. 2020 · 2006
Cited alongside, same era.
First-order definable languages
Volker Diekert and Paul Gastin. 2008 · 2008
Cited alongside, same era.
The dot-depth hierarchy, 45 years later
Jean-Éric Pin. 2017 · 2017
Cited alongside, same era.
Sequential neural networks as automata
William Merrill. 2019 · 2019
Later among the works it cites.
Finite automata can be linearly decoded from language-recognizing RNNs
Joshua J. Michalenko, Ameesh Shah, Abhinav Verma, Swarat Chaudhuri, and Ankit B. Patel. 2019 · 2019
Later among the works it cites.
On the turing completeness of modern neural network architectures
Jorge Pérez, Javier Marinković, and Pablo Barceló. 2019 · 2019
Later among the works it cites.
Visualizing and measuring the geometry of bert
Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim. 2019 · 2019
Later among the works it cites.
On evaluating the generalization of LSTM models in formal languages
Mirac Suzgun, Yonatan Belinkov, and Stuart M. Shieber. 2019b · 2019
Later among the works it cites.
Transformer dissection: An unified understanding for transformer’s attention via the lens of kernel
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
Evaluating the ability of LSTMs to learn context-free grammars
Luzi Sennhauser and Robert Berwick. 2018 · 2018
Cited alongside, same era.
Disan: Directional self-attention network for rnn/cnn-free language understanding
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, Shirui Pan, and C. Zhang. 2018 · 2018
Cited alongside, same era.
Closing brackets with recurrent neural networks
Natalia Skachkova, Thomas Trost, and Dietrich Klakow. 2018 · 2018
Cited alongside, same era.
On the practical computational power of finite precision RNNs for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018 · 2018
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019 · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Later among the works it cites.
Investigating BERT’s knowledge of language: Five analysis methods with NPIs
Alex Warstadt, Yu Cao, Ioana Grosu, Wei Peng, Hagen Blix, Yining Nie, Anna Alsop, Shikha Bordia, Haokun Liu, Alicia Parrish, Sheng-Fu Wang, Jason Phang, Anhad Mohananey, Phu Mon Htut, Paloma Jeretic, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Learning deterministic weighted automata with queries and counterexamples
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2019 · 2019
Later among the works it cites.
Assessing the ability of self-attention networks to learn word order
Baosong Yang, Longyue Wang, Derek F. Wong, Lidia S. Chao, and Zhaopeng Tu. 2019 · 2019
Later among the works it cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Closest in time.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar. 2020 · 2020
Closest in time.