Fetching the paper…
Reading the bibliography…
Machine learning systems perform well on pattern matching tasks, but their ability to perform algorithmic or logical reasoning is not well understood.
Parallel prefix computation
Richard E Ladner and Michael J Fischer · 1980
Earlier work this paper cites.
Lstm recurrent networks learn simple context-free and context-sensitive languages
Felix A Gers and E Schmidhuber · 2001
Earlier work this paper cites.
Tutorial on training recurrent neural networks, covering BPPT, RTRL, EKF and the" echo state network" approach , volume 5.01
Herbert Jaeger · 2002
Earlier work this paper cites.
Training recurrent networks by evolino
Jürgen Schmidhuber, Daan Wierstra, Matteo Gagliolo, and Faustino Gomez · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Self-delimiting neural networks
Jürgen Schmidhuber · 2012
Earlier work this paper cites.
Neural turing machines, 2014
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Łukasz Kaiser and Ilya Sutskever · 2015
Cited alongside, same era.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Cited alongside, same era.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Improving the neural gpu architecture for algorithm learning
Karlis Freivalds and Renars Liepins · 2017
Cited alongside, same era.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2018
Later among the works it cites.
Recurrent relational networks
Rasmus Berg Palm, Ulrich Paquet, and Ole Winther · 2018
Later among the works it cites.
Learning a sat solver from single-bit supervision
Daniel Selsam, Matthew Lamm, Benedikt Bünz, Percy Liang, Leonardo de Moura, and David L Dill · 2018
Later among the works it cites.
Differentiable adaptive computation time for visual reasoning
Cristobal Eyzaguirre and Alvaro Soto · 2020
Later among the works it cites.
Andrea Banino, Jan Balaguer, and Charles Blundell · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Nais-net: Stable deep networks from non-autonomous differential equations
Marco Ciccone, Marco Gallieri, Jonathan Masci, Christian Osendorfer, and Faustino Gomez · 2018
Cited alongside, same era.
Datasets for studying generalization from easy to hard examples, 2021a
Avi Schwarzschild, Eitan Borgnia, Arjun Gupta, Arpit Bansal, Zeyad Emam, Furong Huang, Micah Goldblum, and Tom Goldstein
Cited in the paper.
Can you learn an algorithm? generalizing from easy to hard problems with recurrent networks, 2021b
Avi Schwarzschild, Eitan Borgnia, Arjun Gupta, Furong Huang, Uzi Vishkin, Micah Goldblum, and Tom Goldstein
Cited in the paper.
The uncanny similarity of recurrence and depth, 2021c
Avi Schwarzschild, Arjun Gupta, Amin Ghiasi, Micah Goldblum, and Tom Goldstein
Cited in the paper.
Kwwabena Nuamah · 2021
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A Smith, and Mike Lewis · 2021
Later among the works it cites.