Fetching the paper…
Reading the bibliography…
The following report introduces ideas augmenting standard Long Short Term Memory (LSTM) architecture with multiple memory cells per hidden unit in order to improve its generalization capabilities.
Prediction and entropy of printed english
C. E. Shannon · 1951
Earlier work this paper cites.
Sparse Distributed Memory
P. Kanerva · 1988
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Information Theory, Inference, and Learning Algorithms
D. J. C. MacKay · 2003
Earlier work this paper cites.
Universal Artificial Intelligence: Sequential Decisions based on Algorithmic Probability
M. Hutter · 2005
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Generating text with recurrent neural networks
I. Sutskever, J. Martens, and G. Hinton · 2011
Earlier work this paper cites.
Generating sequences with recurrent neural networks
A. Graves · 2013
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, Ç. Gülçehre, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
How to construct deep recurrent neural networks
R. Pascanu, Ç. Gülçehre, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Recurrent neural network regularization
W. Zaremba, I. Sutskever, and O. Vinyals · 2014
Cited alongside, same era.
Gated feedback recurrent neural networks
J. Chung, Ç. Gülçehre, K. Cho, and Y. Bengio · 2015
Cited alongside, same era.
K. Yao, T. Cohn, K. Vylomova, K. Duh, and C. Dyer · 2015
Later among the works it cites.
T. Cooijmans, N. Ballas, C. Laurent, and A. Courville · 2016
Closest in time.
Persistent rnns: Stashing recurrent weights on-chip
G. Diamos, S. Sengupta, B. Catanzaro, M. Chrzanowski, A. Coates, E. Elsen, J. Engel, A. Hannun, and S. Satheesh · 2016
Closest in time.
Adaptive computation time for recurrent neural networks
A. Graves · 2016
Closest in time.
Zoneout: Regularizing rnns by randomly preserving hidden activations
D. Krueger, T. Maharaj, J. Kramár, M. Pezeshki, N. Ballas, N. R. Ke, A. Goyal, Y. Bengio, H. Larochelle, A. C. Courville, and C. Pal · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Kalchbrenner, I. Danihelka, and A. Graves · 2015
Cited alongside, same era.
Visualizing and understanding recurrent networks
A. Karpathy, J. Johnson, and F. Li · 2015
Cited alongside, same era.
Closest in time.
Recurrent dropout without memory loss
S. Semeniuta, A. Severyn, and E. Barth · 2016
Closest in time.
Recurrent highway networks, 2016
J. G. Zilly, R. K. Srivastava, J. Koutník, and J. Schmidhuber · 2016
Closest in time.