Fetching the paper…
Reading the bibliography…
Transformers are being used extensively across several sequence modeling tasks.
On the computational power of rnns
Samuel A Korsky and Robert C Berwick. 2019 · 1906
Earlier work this paper cites.
A logical calculus of the ideas immanent in nervous activity
Warren S McCulloch and Walter Pitts. 1943 · 1943
Earlier work this paper cites.
On the computational power of neural nets
Hava T Siegelmann and Eduardo D Sontag. 1992 · 1992
Earlier work this paper cites.
A field guide to dynamical recurrent networks
John F Kolen and Stefan C Kremer. 2001 · 2001
Earlier work this paper cites.
Infinite attention: Nngp and ntk for deep attention networks
Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, and Roman Novak. 2020 · 2006
Earlier work this paper cites.
The lipschitz constant of self-attention
Hyunjik Kim, George Papamakarios, and Andriy Mnih. 2020 · 2006
Earlier work this paper cites.
Limits to depth efficiencies of self-attention
Yoav Levine, Noam Wies, Or Sharir, Hofit Bata, and Amnon Shashua. 2020 · 2006
Earlier work this paper cites.
Neural networks and analog computation: beyond the Turing limit
Hava T Siegelmann. 2012 · 2012
Earlier work this paper cites.
Stanford neural machine translation systems for spoken language domains
Minh-Thang Luong and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
OpenNMT: Open-source toolkit for neural machine translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander Rush. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Recurrent neural networks as weighted language recognizers
Yining Chen, Sorcha Gilroy, Andreas Maletti, Jonathan May, and Kevin Knight. 2018 · 2018
Cited alongside, same era.
An improved relative self-attention mechanism for transformer with application to music generation
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Curtis Hawthorne, Andrew M. Dai, Matthew D. Hoffman, and Douglas Eck. 2018 · 2018
Cited alongside, same era.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
On the practical computational power of finite precision RNNs for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018 · 2018
Later among the works it cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
On the turing completeness of modern neural network architectures
Jorge Pérez, Javier Marinković, and Pablo Barceló. 2019 · 2019
Later among the works it cites.
Transformer dissection: An unified understanding for transformer’s attention via the lens of kernel
Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexander Rush. 2018 · 2018
Cited alongside, same era.
Evaluating the ability of LSTMs to learn context-free grammars
Luzi Sennhauser and Robert Berwick. 2018 · 2018
Cited alongside, same era.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Cited alongside, same era.
Disan: Directional self-attention network for rnn/cnn-free language understanding
Tao Shen, Tianyi Zhou, Guodong Long, Jing Jiang, Shirui Pan, and Chengqi Zhang. 2018 · 2018
Cited alongside, same era.
Closing brackets with recurrent neural networks
Natalia Skachkova, Thomas Trost, and Dietrich Klakow. 2018 · 2018
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Later among the works it cites.
Assessing the ability of self-attention networks to learn word order
Baosong Yang, Longyue Wang, Derek F. Wong, Lidia S. Chao, and Zhaopeng Tu. 2019 · 2019
Later among the works it cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Closest in time.
A formal hierarchy of RNN architectures
William Merrill, Gail Weiss, Yoav Goldberg, Roy Schwartz, Noah A. Smith, and Eran Yahav. 2020 · 2020
Closest in time.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar. 2020 · 2020
Closest in time.