Fetching the paper…
Reading the bibliography…
We propose a reparameterization of LSTM that brings the benefits of batch normalization to recurrent neural networks.
Untersuchungen zu dynamischen neuronalen netzen
S. Hochreiter · 1991
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
M. P. Marcus, M. Marcinkiewicz, and B. Santorini · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J Schmidhuber · 1997
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
H. Shimodaira · 2000
Earlier work this paper cites.
Large text compression benchmark
M. Mahoney · 2009
Earlier work this paper cites.
Learning recurrent neural networks with hessian-free optimization
J. Martens and I. Sutskever · 2011
Earlier work this paper cites.
Subword language modeling with neural networks
T. Mikolov, I. Sutskever, A. Deoras, H. Le, S. Kombrink, and J. Cernocky · 2012
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2012
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
A. Graves · 2013
Earlier work this paper cites.
Persistent contextual neural networks for learning symbolic data sequences
Yann Ollivier · 2013
Cited alongside, same era.
Regularization and nonlinearities for neural language models: when are they needed?
Marius Pachitariu and Maneesh Sahani · 2013
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio · 2015
Later among the works it cites.
Describing videos by exploiting temporal structure
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville · 2015
Later among the works it cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Closest in time.
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio · 2016
Closest in time.
David Ha, Andrew Dai, and Quoc V Le · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Amodei et al · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
M. Arjovsky, A. Shah, and Y. Bengio · 2015
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Cited alongside, same era.
Teaching machines to read and comprehend
K. M. Hermann, T. Kocisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, N. Jaitly, and G. Hinton · 2015
Cited alongside, same era.
Blocks and fuel: Frameworks for deep learning
Bart van Merriënboer, Dzmitry Bahdanau, Vincent Dumoulin, Dmitriy Serdyuk, David Warde-Farley, Jan Chorowski, and Yoshua Bengio · 2015
Cited alongside, same era.
Closest in time.
Regularizing rnns by stabilizing activations
D Krueger and R. Memisevic · 2016
Closest in time.
Zoneout: Regularizing rnns by randomly preserving hidden activations
David Krueger, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Hugo Larochelle, and Aaron Courville · 2016
Closest in time.
Batch normalized recurrent neural networks
C. Laurent, G. Pereyra, P. Brakel, Y. Zhang, and Y. Bengio · 2016
Closest in time.
Bridging the gaps between residual learning, recurrent neural networks and visual cortex
Qianli Liao and Tomaso Poggio · 2016
Closest in time.
Theano: A Python framework for fast computation of mathematical expressions
The Theano Development Team et al · 2016
Closest in time.
Architectural complexity measures of recurrent neural networks
S. Zhang, Y. Wu, T. Che, Z. Lin, R. Memisevic, R. Salakhutdinov, and Y. Bengio · 2016
Closest in time.