Fetching the paper…
Reading the bibliography…
We propose Nested LSTMs (NLSTM), a novel RNN architecture with multiple levels of memory.
Untersuchungen zu dynamischen neuronalen netzen
Sepp Hochreiter · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Hierarchical recurrent neural networks for long-term dependencies
Salah El Hihi and Yoshua Bengio · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Learning to forget: Continual prediction with lstm
Felix A. Gers, Jürgen Schmidhuber, and Fred Cummins · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Learning options in reinforcement learning
Martin Stolle and Doina Precup · 2002
Earlier work this paper cites.
Multi-dimensional Recurrent Neural Networks , pages 549–558
Alex Graves, Santiago Fernández, and Jürgen Schmidhuber · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
About the test data, 2011
Matt Mahoney · 2011
Earlier work this paper cites.
Neural networks for machine learning coursera video lectures - geoffrey hinton
Geoffrey Hinton · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Generating Sequences With Recurrent Neural Networks
A. Graves · 2013
Earlier work this paper cites.
Training and analysing deep recurrent neural networks
Michiel Hermans and Benjamin Schrauwen · 2013
Earlier work this paper cites.
How to construct deep recurrent neural networks
Razvan Pascanu, Çaglar Gülçehre, Kyunghyun Cho, and Yoshua Bengio · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons
Aaron Courville Yoshua Bengio, Nicholas Léonard · 2013
Cited alongside, same era.
Pac-inspired option discovery in lifelong reinforcement learning
Emma Brunskill and Lihong Li · 2014
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Çaglar Gülçehre, KyungHyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Blocks and fuel: Frameworks for deep learning
Bart van Merriënboer, Dzmitry Bahdanau, Vincent Dumoulin, Dmitriy Serdyuk, David Warde-Farley, Jan Chorowski, and Yoshua Bengio · 2015
Later among the works it cites.
Reinforcement learning neural turing machines
Wojciech Zaremba and Ilya Sutskever · 2015
Later among the works it cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, et al · 2016
Later among the works it cites.
Classifying options for deep reinforcement learning
Kai Arulkumaran, Nat Dilokthanakul, Murray Shanahan, and Anil Anthony Bharath · 2016
Later among the works it cites.
Using Fast Weights to Attend to the Recent Past
J. Ba, G. Hinton, V. Mnih, J. Z. Leibo, and C. Ionescu · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Cited alongside, same era.
Deep speech: Scaling up end-to-end speech recognition
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Jan Koutnik, Klaus Greff, Faustino Gomez, and Juergen Schmidhuber · 2014
Cited alongside, same era.
Chinese poetry generation with recurrent neural networks
Xingxing Zhang and Mirella Lapata · 2014
Cited alongside, same era.
Gated feedback recurrent neural networks
Junyoung Chung, Caglar Gülçehre, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
Learning to transduce with unbounded memory
Edward Grefenstette, Karl Moritz Hermann, Mustafa Suleyman, and Phil Blunsom · 2015
Cited alongside, same era.
Klaus Greff, Rupesh Kumar Srivastava, Jan Koutník, Bas R. Steunebrink, and Jürgen Schmidhuber · 2015
Cited alongside, same era.
Later among the works it cites.
Using fast weights to attend to the recent past
Jimmy Ba, Geoffrey E Hinton, Volodymyr Mnih, Joel Z Leibo, and Catalin Ionescu · 2016
Later among the works it cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2016
Later among the works it cites.
Long short-term memory-networks for machine reading
Jianpeng Cheng, Li Dong, and Mirella Lapata · 2016
Later among the works it cites.
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio · 2016
Later among the works it cites.
Tim Cooijmans, Nicolas Ballas, César Laurent, Caglar Gulcehre, and Aaron Courville · 2016
Later among the works it cites.
Associative long short-term memory
Ivo Danihelka, Greg Wayne, Benigno Uria, Nal Kalchbrenner, and Alex Graves · 2016
Later among the works it cites.
Zoneout: Regularizing rnns by randomly preserving hidden activations
David Krueger, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Hugo Larochelle, Aaron C. Courville, and Chris Pal · 2016
Later among the works it cites.
Recurrent memory array structures
Kamil Rocki · 2016
Later among the works it cites.
Theano: A Python framework for fast computation of mathematical expressions
Theano Development Team · 2016
Later among the works it cites.
Architectural complexity measures of recurrent neural networks
Saizheng Zhang, Yuhuai Wu, Tong Che, Zhouhan Lin, Roland Memisevic, Ruslan Salakhutdinov, and Yoshua Bengio · 2016
Later among the works it cites.
Maximum-likelihood augmented discrete generative adversarial networks
Tong Che, Yanran Li, Ruixiang Zhang, R Devon Hjelm, Wenjie Li, Yangqiu Song, and Yoshua Bengio · 2017
Later among the works it cites.
Seqgan: sequence generative adversarial nets with policy gradient
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu · 2017
Later among the works it cites.