Fetching the paper…
Reading the bibliography…
We propose a simplified model of attention which is applicable to feed-forward neural networks and demonstrate that the resulting model can solve the synthetic "addition" and "multiplication" long-term memory problems for sequence lengths which are both longer and more widely varying than the best published results for these tasks.
Backpropagation through time: what it does and how to do it
Paul J. Werbos · 1990
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Text categorization with support vector machines: Learning with many relevant features
Thorsten Joachims · 1998
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
James Bergstra, Olivier Breuleux, Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, Guillaume Desjardins, Joseph Turian, David Warde-Farley, and Yoshua Bengio · 2010
Earlier work this paper cites.
Learning recurrent neural networks with hessian-free optimization
James Martens and Ilya Sutskever · 2011
Earlier work this paper cites.
Theano: new features and speed improvements
Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, James Bergstra, Ian Goodfellow, Arnaud Bergeron, Nicolas Bouchard, David Warde-Farley, and Yoshua Bengio · 2012
Earlier work this paper cites.
Long short-term memory in echo state networks: Details of a simulation study
Herbert Jaeger · 2012
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2012
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L. Maas, Awni Y. Hannun, and Andrew Y. Ng · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Recommending music on Spotify with deep learning
Sander Dieleman · 2014
Cited alongside, same era.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Cited alongside, same era.
Introduction to neural machine translation with GPUs (part 3)
Kyunghyun Cho · 2015
Closest in time.
Lasagne: First release
Sander Dieleman, Jan Schlüter, Colin Raffel, Eben Olson, and Soren Kaae Sonderby · 2015
Closest in time.
Regularizing RNNs by stabilizing activations
David Krueger and Roland Memisevic · 2015
Closest in time.
A simple way to initialize recurrent networks of rectified linear units
Quoc V. Le, Navdeep Jaitly, and Geoffrey E. Hinton · 2015
Closest in time.
Molding CNNs for text: non-linear, non-consecutive convolutions
Tao Lei, Regina Barzilay, and Tommi Jaakkola · 2015
Closest in time.
Convolutional lstm networks for subcellular localization of proteins
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2015
Cited alongside, same era.
End-to-end attention-based large vocabulary speech recognition
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio · 2015
Cited alongside, same era.
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals · 2015
Cited alongside, same era.
Søren Kaae Sønderby, Casper Kaae Sønderby, Henrik Nielsen, and Ole Winther · 2015
Closest in time.
Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, and Rob Fergus · 2015
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio · 2015
Closest in time.
Pruning subsequence search with attention-based embedding
Colin Raffel and Daniel P. W. Ellis · 2016
Closest in time.