Fetching the paper…
Reading the bibliography…
Deriving formal bounds on the expressivity of transformers, as well as studying transformers that are constructed to implement known algorithms, are both effective methods for better understanding the computational power of transformers.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
On the temporal analysis of fairness
Dov Gabbay, Amir Pnueli, Saharon Shelah, and Jonathan Stavi · 1990
Earlier work this paper cites.
Tight bounds on the complexity of cascaded decomposition of automata
Oded Maler and Amir Pnueli · 1990
Earlier work this paper cites.
Complexity of modal logics with Presburger constraints
Stéphane Demri and Denis Lugiez · 2010
Earlier work this paper cites.
Hierarchies of piecewise testable languages
Ondřej Klíma and Libor Polák · 2010
Earlier work this paper cites.
An introduction to practical formal methods using temporal logic
Michael Fisher · 2011
Earlier work this paper cites.
Temporal Logic , volume 3
Nicholas Rescher and Alasdair Urquhart · 2012
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Understanding deep neural networks with rectified linear units
Raman Arora, Amitabh Basu, Poorya Mianjy, and Anirbit Mukherjee · 2018
Earlier work this paper cites.
On the ability and limitations of Transformers to recognize formal languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal · 2020
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Cited alongside, same era.
Effects of parameter norm growth during transformer training: Inductive bias from gradient descent
William Merrill, Vivek Ramanujan, Yoav Goldberg, Roy Schwartz, and Noah A. Smith · 2021
Cited alongside, same era.
Attention is Turing-complete
Jorge Pérez, Pablo Barceló, and Javier Marinkovic · 2021
Cited alongside, same era.
Thinking like Transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2021
Cited alongside, same era.
Tracr: Compiled transformers as a laboratory for interpretability
David Lindner, János Kramár, Matthew Rahtz, Thomas McGrath, and Vladimir Mikulik · 2023
Later among the works it cites.
A logic for expressing log-precision transformers
William Merrill and Ashish Sabharwal · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Interleaving logic and counting
Johan van Benthem and Thomas Icard · 2023
Later among the works it cites.
Logical languages accepted by transformer encoders with hard attention
Pablo Barceló, Alexander Kozachinskiy, Anthony Widjaja Lin, and Vladimir Podolskii · 2024
Closest in time.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shunyu Yao, Binghui Peng, Christos Papadimitriou, and Karthik Narasimhan · 2021
Cited alongside, same era.
Masked hard-attention transformers and Boolean RASP recognize exactly the star-free languages, 2023
Dana Angluin, David Chiang, and Andy Yang · 2023
Cited alongside, same era.
On the expressivity role of LayerNorm in Transformers’ attention
Shaked Brody, Uri Alon, and Eran Yahav · 2023
Cited alongside, same era.
Tighter bounds on the expressivity of transformer encoders
David Chiang, Peter Cholak, and Anand Pillay · 2023
Cited alongside, same era.
Closest in time.
The expressive power of transformers with chain of thought
William Merrill and Ashish Sabharwal · 2024
Closest in time.
What formal languages can transformers express? A survey
Lena Strobl, William Merrill, Gail Weiss, David Chiang, and Dana Angluin · 2024
Closest in time.
What algorithms can Transformers learn? A study in length generalization
Hattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin, Omid Saremi, Josh Susskind, Samy Bengio, and Preetum Nakkiran · 2024
Closest in time.