Fetching the paper…
Reading the bibliography…
We propose zoneout, a novel method for regularizing RNNs.
Untersuchungen zu dynamischen neuronalen netzen
Sepp Hochreiter · 1991
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi · 1994
Earlier work this paper cites.
Hierarchical recurrent neural networks for long-term dependencies
Salah El Hihi and Yoshua Bengio · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
Felix A. Gers, Jürgen Schmidhuber, and Fred A. Cummins · 2000
Earlier work this paper cites.
About the test data, 2011
Matt Mahoney · 2011
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Earlier work this paper cites.
Understanding the exploding gradient problem
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
On Fast Dropout and its Applicability to Recurrent Networks
J. Bayer, C. Osendorfer, D. Korhammer, N. Chen, S. Urban, and P. van der Smagt · 2013
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville · 2013
Earlier work this paper cites.
Dropout improves Recurrent Neural Networks for Handwriting Recognition
V. Pham, T. Bluche, C. Kermorvant, and J. Louradour · 2013
Cited alongside, same era.
Fast dropout training
Sida Wang and Christopher Manning · 2013
Cited alongside, same era.
Learning with pseudo-ensembles
Philip Bachman, Ouais Alsharif, and Doina Precup · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Jan Koutnik, Klaus Greff, Faustino Gomez, and Juergen Schmidhuber · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Rnndrop: A novel dropout for rnns in asr
Taesup Moon, Heeyoul Choi, Hoshik Lee, and Inchul Song · 2015
Later among the works it cites.
Blocks and fuel: Frameworks for deep learning
Bart van Merriënboer, Dzmitry Bahdanau, Vincent Dumoulin, Dmitriy Serdyuk, David Warde-Farley, Jan Chorowski, and Yoshua Bengio · 2015
Later among the works it cites.
Lei Jimmy Ba, Ryan Kiros, and Geoffrey E. Hinton · 2016
Closest in time.
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio · 2016
Closest in time.
Tim Cooijmans, Nicolas Ballas, César Laurent, Caglar Gulcehre, and Aaron Courville · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Cited alongside, same era.
Binaryconnect: Training deep neural networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2015
Cited alongside, same era.
A Theoretically Grounded Application of Dropout in Recurrent Neural Networks
Yarin Gal · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Regularizing rnns by stabilizing activations
David Krueger and Roland Memisevic · 2015
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton · 2015
Cited alongside, same era.
Closest in time.
David Ha, Andrew M. Dai, and Quoc V. Le · 2016
Closest in time.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Weinberger · 2016
Closest in time.
Kamil Rocki, Tomasz Kornuta, and Tegan Maharaj · 2016
Closest in time.
Recurrent dropout without memory loss
Stanislau Semeniuta, Aliaksei Severyn, and Erhardt Barth · 2016
Closest in time.
Swapout: Learning an ensemble of deep architectures
S. Singh, D. Hoiem, and D. Forsyth · 2016
Closest in time.
Theano: A Python framework for fast computation of mathematical expressions
Theano Development Team · 2016
Closest in time.