Fetching the paper…
Reading the bibliography…
In this paper, we propose and investigate a novel memory architecture for neural networks called Hierarchical Attentive Memory (HAM).
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 1992
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Morin, Frederic and Bengio, Yoshua · 2005
Earlier work this paper cites.
Handbook of Neuroscience for the Behavioral Sciences
Berntson, G.G. and Cacioppo, J.T · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, Vinod and Hinton, Geoffrey E · 2010
Earlier work this paper cites.
Understanding the exploding gradient problem
Pascanu, Razvan, Mikolov, Tomas, and Bengio, Yoshua · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Earlier work this paper cites.
Graves, Alex, Wayne, Greg, and Danihelka, Ivo · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc VV · 2014
Cited alongside, same era.
Show and tell: A neural image caption generator
Vinyals, Oriol, Toshev, Alexander, Bengio, Samy, and Erhan, Dumitru · 2014
Cited alongside, same era.
Weston, Jason, Chopra, Sumit, and Bordes, Antoine · 2014
Cited alongside, same era.
Learning to transduce with unbounded memory
Grefenstette, Edward, Hermann, Karl Moritz, Suleyman, Mustafa, and Blunsom, Phil · 2015
Cited alongside, same era.
Inferring algorithmic patterns with stack-augmented recurrent nets
Joulin, Armand and Mikolov, Tomas · 2015
Cited alongside, same era.
Gated graph sequence neural networks
Li, Yujia, Tarlow, Daniel, Brockschmidt, Marc, and Zemel, Richard · 2015
Later among the works it cites.
Neural programmer-interpreters
Reed, Scott and de Freitas, Nando · 2015
Later among the works it cites.
Gradient estimation using stochastic computation graphs
Schulman, John, Heess, Nicolas, Weber, Theophane, and Abbeel, Pieter · 2015
Later among the works it cites.
Srivastava, Rupesh Kumar, Greff, Klaus, and Schmidhuber, Jürgen · 2015
Later among the works it cites.
Sukhbaatar, Sainbayar, Szlam, Arthur, Weston, Jason, and Fergus, Rob · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kaiser, Łukasz and Sutskever, Ilya · 2015
Cited alongside, same era.
Kalchbrenner, Nal, Danihelka, Ivo, and Graves, Alex · 2015
Cited alongside, same era.
Kurach, Karol, Andrychowicz, Marcin, and Sutskever, Ilya · 2015
Cited alongside, same era.
Later among the works it cites.
Vinyals, Oriol, Fortunato, Meire, and Jaitly, Navdeep · 2015
Later among the works it cites.
Reinforcement learning neural turing machines
Zaremba, Wojciech and Sutskever, Ilya · 2015
Later among the works it cites.
Learning simple algorithms from examples
Zaremba, Wojciech, Mikolov, Tomas, Joulin, Armand, and Fergus, Rob · 2015
Later among the works it cites.