Fetching the paper…
Reading the bibliography…
What is the computational model behind a Transformer? Where recurrent neural networks have direct parallels in finite state machines, allowing clear discussion and thought around architecture variants or trained models, Transformers have no such familiar parallel.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 1904
Earlier work this paper cites.
What does BERT look at? an analysis of bert’s attention
Clark, K., Khandelwal, U., Levy, O., and Manning, C. D · 1906
Earlier work this paper cites.
Sequential neural networks as automata
Merrill, W · 1906
Earlier work this paper cites.
Finite state automata and simple recurrent networks
Cleeremans, A., Servan-Schreiber, D., and McClelland, J. L · 1989
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Hornik, K., Stinchcombe, M. B., and White, H · 1989
Earlier work this paper cites.
Extraction of rules from discrete-time recurrent neural networks
Omlin, C. W. and Giles, C. L · 1996
Earlier work this paper cites.
ETC: encoding long and structured data in transformers
Ainslie, J., Ontañón, S., Alberti, C., Pham, P., Ravula, A., and Sanghai, S · 2004
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2004
Earlier work this paper cites.
Efficient transformers: A survey
Tay, Y., Dehghani, M., Bahri, D., and Metzler, D · 2009
Earlier work this paper cites.
Parameter norm growth during training of transformers
Merrill, W., Ramanujan, V., Goldberg, Y., Schwartz, R., and Smith, N. A · 2010
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Inferring algorithmic patterns with stack-augmented recurrent nets
Joulin, A. and Mikolov, T · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Luong, T., Pham, H., and Manning, C. D · 2015
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Explaining black boxes on sequential data using weighted automata
Ayache, S., Eyraud, R., and Goudian, N · 2018
Cited alongside, same era.
Can recurrent neural networks learn nested recursion?
Bernardy, J.-P · 2018
Cited alongside, same era.
Evaluating the ability of lstms to learn context-free grammars
Sennhauser, L. and Berwick, R. C · 2018
Transformers as soft reasoners over language
Clark, P., Tafjord, O., and Richardson, K · 2020
Later among the works it cites.
How can self-attention networks recognize Dyck-n languages?
Ebrahimi, J., Gelda, D., and Zhang, W · 2020
Later among the works it cites.
Theoretical limitations of self-attention in neural sequence models
Hahn, M · 2020
Later among the works it cites.
Rnns can generate bounded hierarchical languages with optimal memory
Hewitt, J., Hahn, M., Ganguli, S., Liang, P., and Manning, C. D · 2020
Later among the works it cites.
A formal hierarchy of RNN architectures
Merrill, W., Weiss, G., Goldberg, Y., Schwartz, R., Smith, N. A., and Yahav, E · 2020
Later among the works it cites.
Improving transformer models by reordering their sublayers
Press, O., Smith, N. A., and Levy, O · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Closing brackets with recurrent neural networks
Skachkova, N., Trost, T., and Klakow, D · 2018
Cited alongside, same era.
Universal transformers
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, L · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Connecting weighted automata and recurrent neural networks through spectral learning
Rabusseau, G., Li, T., and Precup, D · 2019
Cited alongside, same era.
On the ability and limitations of transformers to recognize formal languages
Bhattamishra, S., Ahuja, K., and Goyal, N · 2020
Cited alongside, same era.
On the practical computational power of finite precision rnns for language recognition
Weiss, G., Goldberg, Y., and Yahav, E
Cited in the paper.
Are transformers universal approximators of sequence-to-sequence functions?
Yun, C., Bhojanapalli, S., Rawat, A. S., Reddi, S. J., and Kumar, S · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontañón, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A · 2020
Later among the works it cites.
Attention is turing-complete
Pérez, J., Barceló, P., and Marinkovic, J · 2021
Closest in time.
Efficient content-based sparse attention with routing transformers
Roy, A., Saffar, M., Vaswani, A., and Grangier, D · 2021
Closest in time.
Self-attention networks can process bounded hierarchical languages
Yao, S., Peng, B., Papadimitriou, C., and Narasimhan, K · 2021
Closest in time.