Fetching the paper…
Reading the bibliography…
Despite their omnipresence in modern NLP, characterizing the computational power of transformer neural nets remains an interesting open question.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
The circuit value problem is log space complete for P
Richard E Ladner. 1975 · 1975
Earlier work this paper cites.
Complete problems for deterministic polynomial time
Neil D. Jones and William T. Laaser. 1976 · 1976
Earlier work this paper cites.
The complexity of computing the permanent
Leslie G. Valiant. 1979 · 1979
Earlier work this paper cites.
Symmetric space-bounded computation
Harry R. Lewis and Christos H. Papadimitriou. 1982 · 1982
Earlier work this paper cites.
Complexity results for planning
Tom Bylander. 1991 · 1991
Earlier work this paper cites.
A compendium of problems complete for P
Raymond Greenlaw, James M. Hoover, and Walter L. Ruzzo. 1991 · 1991
Earlier work this paper cites.
The permanent requires large uniform threshold circuits
Eric Allender. 1999 · 1999
Earlier work this paper cites.
Division in logspace-uniform nc1
Andrew Chiu, George I. Davida, and Bruce E. Litow. 2001 · 2001
Earlier work this paper cites.
Division is in uniform T C 0 TC^{0}
William Hesse. 2001 · 2001
Cited alongside, same era.
On the linguistic capacity of real-time counter automata
William Cooper Merrill. 2020 · 2004
Cited alongside, same era.
Undirected connectivity in log-space
Omer Reingold. 2008 · 2008
Cited alongside, same era.
Computational Complexity: A Modern Approach
Sanjeev Arora and Boaz Barak. 2009 · 2009
Cited alongside, same era.
Handbook of Satisfiability: Volume 185 Frontiers in Artificial Intelligence and Applications
Arin Biere, Marijn Heule, Hans van Maaren, and Toby Walsh. 2009 · 2009
Cited alongside, same era.
Descriptive complexity
Neil Immerman. 2012 · 2012
Cited alongside, same era.
On the Turing completeness of modern neural network architectures
Jorge Pérez, Javier Marinković, and Pablo Barceló. 2019 · 2019
Later among the works it cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021 · 2021
Later among the works it cites.
Effects of parameter norm growth during transformer training: Inductive bias from gradient descent
William Merrill, Vivek Ramanujan, Yoav Goldberg, Roy Schwartz, and Noah A. Smith. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
GPT3.int8(): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, and Luke Zettlemoyer. 2022 · 2022
Closest in time.
What makes instruction learning hard? An investigation and a new challenge in a synthetic environment
Matthew Finlayson, Kyle Richardson, Ashish Sabharwal, and Peter Clark. 2022 · 2022
Closest in time.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank. 2022 · 2022
Closest in time.
Saturated transformers are constant-depth threshold circuits
William Cooper Merrill, Ashish Sabharwal, and Noah A. Smith. 2022 · 2022
Closest in time.