Fetching the paper…
Reading the bibliography…
We introduce multiplicative LSTM (mLSTM), a recurrent neural network architecture for sequence modelling that combines the long short-term memory (LSTM) and multiplicative recurrent neural network architectures.
Building a large annotated corpus of English: The Penn Treebank
M. P. Marcus, M. A. Marcinkiewicz, and B. Santorini · 1993
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Generating text with recurrent neural networks
I. Sutskever, J. Martens, and G. E. Hinton · 2011
Earlier work this paper cites.
The human knowledge compression contest
M. Hutter · 2012
Earlier work this paper cites.
Subword language modeling with neural networks
T. Mikolov, I. Sutskever, A. Deoras, H. Le, S. Kombrink, and J. Cernocky · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. E. Hinton · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
A. Graves · 2013
Earlier work this paper cites.
Regularization and nonlinearities for neural language models: when are they needed?
M. Pachitariu and M. Sahani · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
Neural machine translation in linear time
N. Kalchbrenner, L. Espeholt, K. Simonyan, A. Oord, A. Graves, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier · 2017
Closest in time.
Hypernetworks
D. Ha, A. Dai, and Q. Lee · 2017
Closest in time.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher · 2017
Closest in time.
Dynamic evaluation of neural sequence models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals · 2017
Closest in time.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Diederik P Kingma · 2016
Cited alongside, same era.
On multiplicative integration with recurrent neural networks
Yuhuai Wu, Saizheng Zhang, Ying Zhang, Yoshua Bengio, and Ruslan R Salakhutdinov · 2016
Cited alongside, same era.
Architectural complexity measures of recurrent neural networks
S. Zhang, Y. Wu, T. Che, Z. Lin, R. Memisevic, R. Salakhutdinov, and Y. Bengio · 2016
Cited alongside, same era.
Hierarchical multiscale recurrent neural networks
J. Chung, S. Ahn, and Y. Bengio · 2017
Cited alongside, same era.
Recurrent batch normalization
Tim Cooijmans, Nicolas Ballas, César Laurent, and Aaron Courville · 2017
Cited alongside, same era.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher
Cited in the paper.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher
Cited in the paper.
Asier Mujika, Florian Meier, and Angelika Steger · 2017
Closest in time.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2017
Closest in time.
Learning to generate reviews and discovering sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever · 2017
Closest in time.
Recurrent highway networks
J. G. Zilly, R. K. Srivastava, J. Koutník, and J. Schmidhuber · 2017
Closest in time.