Fetching the paper…
Reading the bibliography…
Despite the widespread success of Transformers on NLP tasks, recent works have found that they struggle to model several formal languages when compared to recurrent models.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
On the computational power of rnns
Samuel A Korsky and Robert C Berwick. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Neural networks are a priori biased towards boolean functions with low entropy
Chris Mingard, Joar Skalse, Guillermo Valle-Pérez, David Martínez-Rubio, Vladimir Mikulik, and Ard A Louis. 2019 · 1909
Earlier work this paper cites.
Random deep neural networks are biased towards simple functions
Giacomo De Palma, Bobak Toussi Kiani, and Seth Lloyd. 2019 · 1974
Earlier work this paper cites.
The influence of variables on Boolean functions
Jeff Kahn, Gil Kalai, and Nathan Linial. 1989 · 1989
Earlier work this paper cites.
First-order versus second-order single-layer recurrent neural networks
Mark W Goudreau, C Lee Giles, Srimat T Chakradhar, and Dong Chen. 1994 · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael Kearns. 1998 · 1998
Earlier work this paper cites.
Lstm recurrent networks learn simple context-free and context-sensitive languages
Felix A Gers and E Schmidhuber. 2001 · 2001
Earlier work this paper cites.
A field guide to dynamical recurrent networks
John F Kolen and Stefan C Kremer. 2001 · 2001
Earlier work this paper cites.
A formal hierarchy of rnn architectures
William Merrill, Gail Weiss, Yoav Goldberg, Roy Schwartz, Noah A Smith, and Eran Yahav. 2020 · 2004
Earlier work this paper cites.
Generalization ability of boolean functions implemented in feedforward neural networks
Leonardo Franco. 2006 · 2006
Earlier work this paper cites.
How can self-attention networks recognize dyck-n languages?
Javid Ebrahimi, Dhruv Gelda, and Wei Zhang. 2020 · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Tighter relations between sensitivity and other complexity measures
Andris Ambainis, Mohammad Bavarian, Yihan Gao, Jieming Mao, Xiaoming Sun, and Song Zuo. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Smooth boolean functions are easy: Efficient algorithms for low-sensitivity functions
Parikshit Gopalan, Noam Nisan, Rocco A Servedio, Kunal Talwar, and Avi Wigderson. 2016 · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzundefinedbski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S. Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien. 2017 · 2017
Cited alongside, same era.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein. 2017 · 2017
Cited alongside, same era.
Sympy: symbolic computing in python
Aaron Meurer, Christopher P. Smith, Mateusz Paprocki, Ondřej Čertík, Sergey B. Kirpichev, Matthew Rocklin, AMiT Kumar, Sergiu Ivanov, Jason K. Moore, Sartaj Singh, Thilina Rathnayake, Sean Vig, Brian E. Granger, Richard P. Muller, Francesco Bonazzi, Harsh Gupta, Shivam Vats, Fredrik Johansson, Fabian Pedregosa, Matthew J. Curry, Andy R. Terrel, Štěpán Roučka, Ashutosh Saboo, Isuru Fernando, Sumith Kulal, Robert Cimrman, and Anthony Scopatz. 2017 · 2017
Cited alongside, same era.
Attention is all you need
On the Ability and Limitations of Transformers to Recognize Formal Languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal. 2020a · 2020
Later among the works it cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Later among the works it cites.
Bayesian deep learning and a probabilistic perspective of generalization
Andrew G Wilson and Pavel Izmailov. 2020 · 2020
Later among the works it cites.
Sensitivity as a complexity measure for sequence classification tasks
Michael Hahn, Dan Jurafsky, and Richard Futrell. 2021 · 2021
Later among the works it cites.
Is sgd a bayesian sampler? well, almost
Chris Mingard, Guillermo Valle-Pérez, Joar Skalse, and Ard A Louis. 2021 · 2021
Later among the works it cites.
Analysis of boolean functions
Ryan O’Donnell. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar. 2018 · 2018
Cited alongside, same era.
Sensitivity and generalization in neural networks: an empirical study
Roman Novak, Yasaman Bahri, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein. 2018 · 2018
Cited alongside, same era.
The annotated transformer
Alexander Rush. 2018 · 2018
Cited alongside, same era.
Evaluating the ability of LSTMs to learn context-free grammars
Luzi Sennhauser and Robert Berwick. 2018 · 2018
Cited alongside, same era.
Closing brackets with recurrent neural networks
Natalia Skachkova, Thomas Trost, and Dietrich Klakow. 2018 · 2018
Cited alongside, same era.
A comparative study of rule extraction for recurrent neural networks
Qinglong Wang, Kaixuan Zhang, Alexander G. Ororbia II, Xinyu Xing, Xue Liu, and C. Lee Giles. 2018 · 2018
Cited alongside, same era.
On the practical computational power of finite precision RNNs for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018 · 2018
Cited alongside, same era.
Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks
Avi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang, Uzi Vishkin, Micah Goldblum, and Tom Goldstein. 2021 · 2021
Later among the works it cites.
Evaluating Transformer’s Ability to Learn Mildly Context-Sensitive Languages
Shunjie Wang. 2021 · 2021
Later among the works it cites.
Self-attention networks can process bounded hierarchical languages
Shunyu Yao, Binghui Peng, Christos Papadimitriou, and Karthik Narasimhan. 2021 · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. 2021 · 2021
Later among the works it cites.
Exploring length generalization in large language models
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur. 2022 · 2022
Closest in time.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Boaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang. 2022 · 2022
Closest in time.
Overcoming a theoretical limitation of self-attention
David Chiang and Peter Cholak. 2022 · 2022
Closest in time.
Neural networks and the chomsky hierarchy
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Marcus Hutter, Shane Legg, and Pedro A Ortega. 2022 · 2022
Closest in time.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang. 2022 · 2022
Closest in time.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank. 2022 · 2022
Closest in time.
Saturated transformers are constant-depth threshold circuits
William Merrill, Ashish Sabharwal, and Noah A. Smith. 2022 · 2022
Closest in time.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra. 2022 · 2022
Closest in time.
Memorisation versus generalisation in pre-trained language models
Michael Tänzer, Sebastian Ruder, and Marek Rei. 2022 · 2022
Closest in time.