Fetching the paper…
Reading the bibliography…
Transformers have become a standard neural network architecture for many NLP problems, motivating theoretical analysis of their power in terms of formal languages.
A history of the prime number theorem
Larry J Goldstein. 1973 · 1973
Earlier work this paper cites.
Parity, circuits, and the polynomial-time hierarchy
Merrick Furst, James B. Saxe, and Michael Sipser. 1981 · 1981
Earlier work this paper cites.
On ACC and threshold circuits
Andrew C.-C. Yao. 1990 · 1990
Earlier work this paper cites.
Machine Models and Simulations , chapter 1. MIT Press, Cambridge, MA, USA
Peter van Emde Boas. 1991 · 1991
Earlier work this paper cites.
A Catalog of Complexity Classes , chapter 2. MIT Press, Cambridge, MA, USA
David S. Johnson. 1991 · 1991
Earlier work this paper cites.
Polynomial size log depth circuits: Between 𝖭𝖢 1 \mathsf{NC}^{1} and 𝖠𝖢 1 \mathsf{AC}^{1}
Meena Mahajan. 2007 · 2007
Earlier work this paper cites.
Computational Complexity: A Modern Approach
Sanjeev Arora and Boaz Barak. 2009 · 2009
Earlier work this paper cites.
RNNs can generate bounded hierarchical languages with optimal memory
John Hewitt, Michael Hahn, Surya Ganguli, Percy Liang, and Christopher D. Manning. 2020 · 2010
Earlier work this paper cites.
Lecture notes for topics in complexity theory
Neeraj Kayal. 2015 · 2015
Earlier work this paper cites.
Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Cited alongside, same era.
Advanced Studies on the Complexity of Formal Languages
Liliana Cojocaru. 2016 · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Rational recurrences
Hao Peng, Roy Schwartz, Sam Thomson, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
On the practical computational power of finite precision RNNs for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
On the ability and limitations of transformers to recognize formal languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal. 2020 · 2020
Later among the works it cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Later among the works it cites.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar. 2020 · 2020
Later among the works it cites.
Formal language theory meets modern NLP
William Merrill. 2021 · 2021
Closest in time.
Effects of parameter norm growth during transformer training: Inductive bias from gradient descent
William Merrill, Vivek Ramanujan, Yoav Goldberg, Roy Schwartz, and Noah A. Smith. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sequential neural networks as automata
William Merrill. 2019 · 2019
Cited alongside, same era.
On the Turing completeness of modern neural network architectures
Jorge Pérez, Javier Marinković, and Pablo Barceló. 2019 · 2019
Cited alongside, same era.
Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP . Association for Computational Linguistics, Online
Afra Alishahi, Yonatan Belinkov, Grzegorz Chrupała, Dieuwke Hupkes, Yuval Pinter, and Hassan Sajjad, editors. 2020 · 2020
Cited alongside, same era.
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2021 · 2021
Closest in time.
Self-attention networks can process bounded hierarchical languages
Shunyu Yao, Binghui Peng, Christos Papadimitriou, and Karthik Narasimhan. 2021 · 2021
Closest in time.
Hard attention transformers and constant depth circuits
Yiding Hao, Dana Angluin, and Robert Frank. 2022 · 2022
Closest in time.