Fetching the paper…
Reading the bibliography…
Existing work has analyzed the representational capacity of the transformer architecture by means of formal models of computation.
A mathematical theory of communication
C. E. Shannon. 1948 · 1948
Earlier work this paper cites.
Neural Nets and the Brain Model Problem
Marvin Lee Minsky. 1954 · 1954
Earlier work this paper cites.
Representation of events in nerve nets and finite automata
S. C. Kleene. 1956 · 1956
Earlier work this paper cites.
Formal language theory: Refining the Chomsky hierarchy
Gerhard Jäger and James Rogers. 2012 · 1970
Earlier work this paper cites.
Continuous speech recognition by statistical methods
F. Jelinek. 1976 · 1976
Earlier work this paper cites.
A maximum likelihood approach to continuous speech recognition
Lalit R. Bahl, Frederick Jelinek, and Robert L. Mercer. 1983 · 1983
Earlier work this paper cites.
Self-Organized Language Modeling for Speech Recognition , page 450–506. Morgan Kaufmann Publishers Inc., San Francisco, CA, USA
F. Jelinek. 1990 · 1990
Earlier work this paper cites.
On the computational power of sigmoid versus boolean threshold circuits
W. Maass, G. Schnitger, and E. D. Sontag. 1991 · 1991
Earlier work this paper cites.
On the computational power of neural nets
Hava T. Siegelmann and E. D. Sontag. 1992 · 1992
Earlier work this paper cites.
Mathematical Perspectives on Neural Networks
Paul Smolensky, Michael C. Mozer, and David E. Rumelhart. 1996 · 1996
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent. 2000 · 2000
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin. 2003 · 2003
Earlier work this paper cites.
Neural Probabilistic Language Models , pages 137–186. Springer Berlin Heidelberg, Berlin, Heidelberg
Yoshua Bengio, Holger Schwenk, Jean-Sébastien Senécal, Fréderic Morin, and Jean-Luc Gauvain. 2006 · 2006
Earlier work this paper cites.
Continuous space language models
Holger Schwenk. 2007 · 2007
Earlier work this paper cites.
RNNs can generate bounded hierarchical languages with optimal memory
John Hewitt, Michael Hahn, Surya Ganguli, Percy Liang, and Christopher D. Manning. 2020 · 2010
Earlier work this paper cites.
KenLM: Faster and smaller language model queries
Kenneth Heafield. 2011 · 2011
Earlier work this paper cites.
Scalable modified Kneser-Ney language model estimation
Kenneth Heafield, Ivan Pouzyrevsky, Jonathan H. Clark, and Philipp Koehn. 2013 · 2013
Cited alongside, same era.
Learning subregular classes of languages with factored deterministic automata
Jeffrey Heinz and James Rogers. 2013 · 2013
Cited alongside, same era.
From softmax to sparsemax: A sparse model of attention and multi-label classification
André F. T. Martins and Ramón F. Astudillo. 2016 · 2016
Cited alongside, same era.
Subregular complexity and deep learning
Enes Avcu, Chihiro Shibata, and Jeffrey Heinz. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Context-free transductions with neural stacks
Extracting finite automata from RNNs using state merging
William Merrill and Nikolaos Tsilivis. 2022 · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2022 · 2022
Later among the works it cites.
An explanation of in-context learning as implicit bayesian inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2022 · 2022
Later among the works it cites.
Masked hard-attention transformers and boolean rasp recognize exactly the star-free languages
Dana Angluin, David Chiang, and Andy Yang. 2023 · 2023
Later among the works it cites.
Tighter bounds on the expressivity of transformer encoders
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yiding Hao, William Merrill, Dana Angluin, Robert Frank, Noah Amsel, Andrew Benz, and Simon Mendelsohn. 2018 · 2018
Cited alongside, same era.
Breaking the softmax bottleneck: A high-rank RNN language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen. 2018 · 2018
Cited alongside, same era.
Sequential neural networks as automata
William Merrill. 2019 · 2019
Cited alongside, same era.
On the ability and limitations of transformers to recognize formal languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal. 2020 · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Cited alongside, same era.
A formal hierarchy of RNN architectures
William Merrill, Gail Weiss, Yoav Goldberg, Roy Schwartz, Noah A. Smith, and Eran Yahav. 2020 · 2020
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021 · 2021
Cited alongside, same era.
David Chiang, Peter Cholak, and Anand Pillay. 2023 · 2023
Later among the works it cites.
A measure-theoretic characterization of tight language models
Li Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, and Ryan Cotterell. 2023 · 2023
Later among the works it cites.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. 2023 · 2023
Later among the works it cites.
On the representational capacity of recurrent neural language models
Franz Nowak, Anej Svete, Li Du, and Ryan Cotterell. 2023 · 2023
Later among the works it cites.
Transformers as recognizers of formal languages: A survey on expressivity
Lena Strobl, William Merrill, Gail Weiss, David Chiang, and Dana Angluin. 2023 · 2023
Later among the works it cites.
Recurrent neural language models as probabilistic finite-state automata
Anej Svete and Ryan Cotterell. 2023 · 2023
Later among the works it cites.
Transformers are uninterpretable with myopic methods: A case study with bounded Dyck grammars
Kaiyue Wen, Yuchen Li, Bingbin Liu, and Andrej Risteski. 2023 · 2023
Later among the works it cites.
Self-attention networks can process bounded hierarchical languages
Shunyu Yao, Binghui Peng, Christos Papadimitriou, and Karthik Narasimhan. 2021 · 2023
Later among the works it cites.
In-context language learning: Architectures and algorithms
Ekin Akyürek, Bailin Wang, Yoon Kim, and Jacob Andreas. 2024 · 2024
Closest in time.
On the representational capacity of neural language models with chain-of-thought reasoning
Franz Nowak, Anej Svete, Alexandra Butoi, and Ryan Cotterell. 2024 · 2024
Closest in time.
A theoretical result on the inductive bias of RNN language models
Anej Svete, Robin Shing Moon Chan, and Ryan Cotterell. 2024 · 2024
Closest in time.