Fetching the paper…
Reading the bibliography…
Reliable generalization lies at the heart of safe ML and AI.
Three models for the description of language
Noam Chomsky · 1956
Earlier work this paper cites.
Computation: Finite and Infinite Machines
Marvin L. Minsky · 1967
Earlier work this paper cites.
The need for biases in learning generalizations
Tom M. Mitchell · 1980
Earlier work this paper cites.
Present position and potential developments: Some personal views: Statistical theory: The prequential approach
A. Philip Dawid · 1984
Earlier work this paper cites.
Genetic Algorithms in Search Optimization and Machine Learning
David E. Goldberg · 1989
Earlier work this paper cites.
Finding structure in time
Jeffrey L. Elman · 1990
Earlier work this paper cites.
The induction of dynamical recognizers
Jordan B. Pollack · 1991
Earlier work this paper cites.
Learning and extracting finite state automata with second-order recurrent neural networks
C. Lee Giles, Clifford B. Miller, Dong Chen, Hsing-Hen Chen, Guo-Zheng Sun, and Yee-Chun Lee · 1992
Earlier work this paper cites.
Adaptation in Natural and Artificial Systems: An Introductory Analysis with Applications to Biology, Control, and Artificial Intelligence
John H. Holland · 1992
Earlier work this paper cites.
A connectionist symbol manipulator that discovers the structure of context-free languages
Michael Mozer and Sreerupa Das · 1992
Earlier work this paper cites.
The neural network pushdown automaton: model, stack and learning simulations
G. Z. Sun, C. L. Giles, H. H. Chen, and Y. C. Lee · 1993
Earlier work this paper cites.
Analog computation via neural networks
Hava T. Siegelmann and Eduardo D. Sontag · 1994
Earlier work this paper cites.
A representation scheme to perform program induction in a canonical genetic algorithm
Mark Wineberg and Franz Oppacher · 1994
Earlier work this paper cites.
Learning to count without a counter: A case study of dynamics and activation landscapes in recurrent networks
Janet Wiles and Jeff Elman · 1995
Earlier work this paper cites.
A recurrent network that performs a context-sensitive prediction task
Mark Steijvers and Peter Grünwald · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Designing a counter: Another case study of dynamics and activation landscapes in recurrent networks
Steffen Hölldobler, Yvonne Kalinke, and Helko Lehmann · 1997
Earlier work this paper cites.
Evolutionary program induction of binary machine code and its applications
Peter Nordin · 1997
Earlier work this paper cites.
Recurrent neural networks can learn to implement symbol-sensitive counting
Paul Rodriguez and Janet Wiles · 1997
Earlier work this paper cites.
Introduction to the theory of computation
Michael Sipser · 1997
Earlier work this paper cites.
Models of computation - exploring the power of computing
John E. Savage · 1998
Earlier work this paper cites.
Statistical learning theory
Vladimir Vapnik · 1998
Earlier work this paper cites.
Context-free and context-sensitive dynamics in recurrent neural networks
Mikael Bodén and Janet Wiles · 2000
Earlier work this paper cites.
LSTM recurrent networks learn simple context-free and context-sensitive languages
Felix A. Gers and Jürgen Schmidhuber · 2001
Earlier work this paper cites.
On learning context-free and context-sensitive languages
Mikael Bodén and Janet Wiles · 2002
Earlier work this paper cites.
Automata, Computability and Complexity
Elaine Rich · 2007
Earlier work this paper cites.
Accelerated neural evolution through cooperatively coevolved synapses
Faustino J. Gomez, Jürgen Schmidhuber, and Risto Miikkulainen · 2008
Earlier work this paper cites.
Algorithmic probability: Theory and applications
Ray J. Solomonoff · 2009
Earlier work this paper cites.
Algorithmic probability, heuristic programming and agi
Ray J. Solomonoff · 2010
Earlier work this paper cites.
Learning dependency-based compositional semantics
Percy Liang, Michael I. Jordan, and Dan Klein · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Cited alongside, same era.
Neural turing machines
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
Learning to transduce with unbounded memory
Edward Grefenstette, Karl Moritz Hermann, Mustafa Suleyman, and Phil Blunsom · 2015
Cited alongside, same era.
Inferring algorithmic patterns with stack-augmented recurrent nets
Armand Joulin and Tomás Mikolov · 2015
Memory architectures in recurrent neural network language models
Dani Yogatama, Yishu Miao, Gábor Melis, Wang Ling, Adhiguna Kuncoro, Chris Dyer, and Phil Blunsom · 2018
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc Viet Le, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
On the computational power of rnns
Samuel A. Korsky and Robert C. Berwick · 2019
Later among the works it cites.
Sequential neural networks as automata
William Merrill · 2019
Later among the works it cites.
On the turing completeness of modern neural network architectures
Jorge Pérez, Javier Marinkovic, and Pablo Barceló · 2019
Later among the works it cites.
A survey of neural networks and formal languages
Joshua Ackerman and George Cybenko · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
End-to-end memory networks
Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, and Rob Fergus · 2015
Cited alongside, same era.
Memory networks
Jason Weston, Sumit Chopra, and Antoine Bordes · 2015
Cited alongside, same era.
Reinforcement learning neural turing machines
Wojciech Zaremba and Ilya Sutskever · 2015
Cited alongside, same era.
Associative long short-term memory
Ivo Danihelka, Greg Wayne, Benigno Uria, Nal Kalchbrenner, and Alex Graves · 2016
Cited alongside, same era.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Cited alongside, same era.
The DeepMind JAX Ecosystem, 2020
Igor Babuschkin, Kate Baumli, Alison Bell, Surya Bhupatiraju, Jake Bruce, Peter Buchlovsky, David Budden, Trevor Cai, Aidan Clark, Ivo Danihelka, Claudio Fantacci, Jonathan Godwin, Chris Jones, Tom Hennigan, Matteo Hessel, Steven Kapturowski, Thomas Keck, Iurii Kemaev, Michael King, Lena Martens, Vladimir Mikulik, Tamara Norman, John Quan, George Papamakarios, Roman Ring, Francisco Ruiz, Alvaro Sanchez, Rosalia Schneider, Eren Sezener, Stephen Spencer, Srivatsan Srinivasan, Wojciech Stokowiec, and Fabio Viola · 2020
Later among the works it cites.
On the ability and limitations of transformers to recognize formal languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal · 2020
Later among the works it cites.
Learning context-free languages with nondeterministic stack rnns
Brian DuSell and David Chiang · 2020
Later among the works it cites.
How can self-attention networks recognize dyck-n languages?
Javid Ebrahimi, Dhruv Gelda, and Wei Zhang · 2020
Later among the works it cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Later among the works it cites.
Haiku: Sonnet for JAX, 2020
Tom Hennigan, Trevor Cai, Tamara Norman, and Igor Babuschkin · 2020
Later among the works it cites.
Optax: composable gradient transformation and optimisation, in jax!, 2020
Matteo Hessel, David Budden, Fabio Viola, Mihaela Rosca, Eren Sezener, and Tom Hennigan · 2020
Later among the works it cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
A formal hierarchy of RNN architectures
William Merrill, Gail Weiss, Yoav Goldberg, Roy Schwartz, Noah A. Smith, and Eran Yahav · 2020
Later among the works it cites.
Provably stable interpretable encodings of context free grammars in rnns with a differentiable stack
John Stogin, Ankur Arjun Mali, and C. Lee Giles · 2020
Later among the works it cites.
Recognizing long grammatical sequences using recurrent networks augmented with an external differentiable stack
Ankur Arjun Mali, Alexander Ororbia, Daniel Kifer, and C. Lee Giles · 2021
Later among the works it cites.
Attention is turing-complete
Jorge Pérez, Pablo Barceló, and Javier Marinkovic · 2021
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A. Smith, and Mike Lewis · 2021
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding, 2021
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Later among the works it cites.
Thinking like transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2021
Later among the works it cites.
Exploring length generalization in large language models
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay V. Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur · 2022
Closest in time.
End-to-end algorithm synthesis with recurrent networks: Logical extrapolation without overthinking
Arpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam, Furong Huang, Micah Goldblum, and Tom Goldstein · 2022
Closest in time.
Overcoming a theoretical limitation of self-attention, 2022
David Chiang and Peter Cholak · 2022
Closest in time.
Learning hierarchical structures with differentiable nondeterministic stacks
Brian DuSell and David Chiang · 2022
Closest in time.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank · 2022
Closest in time.
Log-precision transformers are constant-depth uniform threshold circuits
William Merrill and Ashish Sabharwal · 2022
Closest in time.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harrison Edwards, Igor Babuschkin, and Vedant Misra · 2022
Closest in time.
Unveiling transformers with LEGO: a synthetic reasoning task
Yi Zhang, Arturs Backurs, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, and Tal Wagner · 2022
Closest in time.