Fetching the paper…
Reading the bibliography…
This paper introduces Adaptive Computation Time (ACT), an algorithm that allows recurrent neural networks to learn how many computational steps to take between receiving an input and emitting an output.
Paradigms and processes in reading comprehension
M. A. Just, P. A. Carpenter, and J. D. Woolley · 1982
Earlier work this paper cites.
Modular elliptic curves and fermat’s last theorem
A. J. Wiles · 1995
Earlier work this paper cites.
Gradient-based learning algorithms for recurrent networks and their computational complexity
R. J. Williams and D. Zipser · 1995
Earlier work this paper cites.
The effects of adding noise during backpropagation training on a generalization performance
G. An · 1996
Earlier work this paper cites.
Emergence of simple-cell receptive field properties by learning a sparse code for natural images
B. A. Olshausen et al · 1996
Earlier work this paper cites.
Guessing can outperform many long time lag algorithms
J. Schmidhuber and S. Hochreiter · 1996
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001
S. Hochreiter, Y. Bengio, P. Frasconi, and J. Schmidhuber · 2001
Earlier work this paper cites.
Universal artificial intelligence
M. Hutter · 2005
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Multi-column deep neural networks for image classification
D. C. Ciresan, U. Meier, and J. Schmidhuber · 2012
Earlier work this paper cites.
Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition
G. Dahl, D. Yu, L. Deng, and A. Acero · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Self-delimiting neural networks
J. Schmidhuber · 2012
Cited alongside, same era.
Generating sequences with recurrent neural networks
A. Graves · 2013
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
A. Graves, A. Mohamed, and G. Hinton · 2013
Cited alongside, same era.
An introduction to Kolmogorov complexity and its applications
M. Li and P. Vitányi · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Later among the works it cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Later among the works it cites.
Conditional computation in neural networks for faster models
E. Bengio, P.-L. Bacon, J. Pineau, and D. Precup · 2015
Later among the works it cites.
Learning to transduce with unbounded memory
E. Grefenstette, K. M. Hermann, M. Suleyman, and P. Blunsom · 2015
Later among the works it cites.
Draw: A recurrent neural network for image generation
K. Gregor, I. Danihelka, A. Graves, and D. Wierstra · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Cited alongside, same era.
First experiments with powerplay
R. K. Srivastava, B. R. Steunebrink, and J. Schmidhuber · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Deep sequential neural network
L. Denoyer and P. Gallinari · 2014
Cited alongside, same era.
A. Graves, G. Wayne, and I. Danihelka · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Distributed representations of sentences and documents
Q. V. Le and T. Mikolov · 2014
Cited alongside, same era.
N. Kalchbrenner, I. Danihelka, and A. Graves · 2015
Later among the works it cites.
Neural programmer-interpreters
S. Reed and N. de Freitas · 2015
Later among the works it cites.
Training very deep networks
R. K. Srivastava, K. Greff, and J. Schmidhuber · 2015
Later among the works it cites.
End-to-end memory networks
S. Sukhbaatar, J. Weston, R. Fergus, et al · 2015
Later among the works it cites.
Order matters: Sequence to sequence for sets
O. Vinyals, S. Bengio, and M. Kudlur · 2015
Later among the works it cites.
Pointer networks
O. Vinyals, M. Fortunato, and N. Jaitly · 2015
Later among the works it cites.
Attend, infer, repeat: Fast scene understanding with generative models
S. Eslami, N. Heess, T. Weber, Y. Tassa, K. Kavukcuoglu, and G. E. Hinton · 2016
Closest in time.