Fetching the paper…
Reading the bibliography…
Recurrent neural networks have gained widespread use in modeling sequential data.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1986
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman · 1990
Earlier work this paper cites.
Numerical solution of boundary value problems for ordinary differential equations , volume 13
Uri M Ascher, Robert MM Mattheij, and Robert D Russell · 1994
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Hierarchical recurrent neural networks for long-term dependencies
Salah El Hihi and Yoshua Bengio · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Computer methods for ordinary differential equations and differential-algebraic equations , volume 61
Uri M Ascher and Linda R Petzold · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Learning long term dependencies via Fourier recurrent units
Jiong Zhang, Yibo Lin, Zhao Song, and Inderjit Dhillon · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Parsing natural scenes and natural language with recursive neural networks
Richard Socher, Cliff C Lin, Chris Manning, and Andrew Y Ng · 2011
Earlier work this paper cites.
Understanding the exploding gradient problem
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Cited alongside, same era.
Learning longer memory in recurrent neural networks
Tomas Mikolov, Armand Joulin, Sumit Chopra, Michael Mathieu, and Marc’Aurelio Ranzato · 2014
Cited alongside, same era.
Session-based recommendations with recurrent neural networks
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Later among the works it cites.
Full-capacity unitary recurrent neural networks
Scott Wisdom, Thomas Powers, John Hershey, Jonathan Le Roux, and Les Atlas · 2016
Later among the works it cites.
Recurrent batch normalization
Tim Cooijmans, Nicolas Ballas, César Laurent, Çağlar Gülçehre, and Aaron Courville · 2017
Later among the works it cites.
Stable architectures for deep neural networks
Eldad Haber and Lars Ruthotto · 2017
Later among the works it cites.
Learning unitary operators with help from u (n)
Stephanie L Hyland and Gunnar Rätsch · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
An empirical exploration of recurrent network architectures
Rafal Jozefowicz, Wojciech Zaremba, and Ilya Sutskever · 2015
Cited alongside, same era.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Ruslan R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton · 2015
Cited alongside, same era.
Dmytro Mishkin and Jiri Matas · 2015
Cited alongside, same era.
Multi-scale context aggregation by dilated convolutions
Fisher Yu and Vladlen Koltun · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Cited alongside, same era.
Cijo Jose, Moustpaha Cisse, and Francois Fleuret · 2017
Later among the works it cites.
Preventing gradient explosions in gated recurrent units
Sekitoshi Kanai, Yasuhiro Fujiwara, and Sotetsu Iwamura · 2017
Later among the works it cites.
A recurrent neural network without chaos
Thomas Laurent and James von Brecht · 2017
Later among the works it cites.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Sam Schoenholz, and Surya Ganguli · 2017
Later among the works it cites.
On orthogonality and learning recurrent networks with long term dependencies
Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal · 2017
Later among the works it cites.
Recurrent recommender networks
Chao-Yuan Wu, Amr Ahmed, Alex Beutel, Alexander J Smola, and How Jing · 2017
Later among the works it cites.
Di Xie, Jiang Xiong, and Shiliang Pu · 2017
Later among the works it cites.
Beyond finite layer neural networks: Bridging deep architectures and numerical differential equations
Yiping Lu, Aoxiao Zhong, Quanzheng Li, and Bin Dong · 2018
Later among the works it cites.
Can recurrent neural networks warp time?
Corentin Tallec and Yann Ollivier · 2018
Later among the works it cites.
Residual recurrent neural networks for learning sequential representations
Boxuan Yue, Junwei Fu, and Jun Liang · 2018
Later among the works it cites.