Fetching the paper…
Reading the bibliography…
We present methodology for using dynamic evaluation to improve neural sequence models.
Speech recognition and the frequency of recently used words: A modified Markov model for natural language
R. Kuhn · 1988
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
P. J. Werbos · 1990
Earlier work this paper cites.
A dynamic language model for speech recognition
F. Jelinek, B. Merialdo, S. Roukos, and M. Strauss · 1991
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
M. P. Marcus, M. A. Marcinkiewicz, and B. Santorini · 1993
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Statistical language model adaptation: review and perspectives
J. R. Bellegarda · 2004
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
P. Koehn · 2005
Earlier work this paper cites.
The human knowledge compression prize
M. Hutter · 2006
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
Context dependent recurrent neural network language model
T. Mikolov and G. Zweig · 2012
Earlier work this paper cites.
Subword language modeling with neural networks
T. Mikolov, I. Sutskever, A. Deoras, H. Le, S. Kombrink, and J. Cernocky · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. E. Hinton · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
A. Graves · 2013
Cited alongside, same era.
Regularization of neural networks using dropconnect
L. Wan, M. Zeiler, S. Zhang, Yann L. Cun, and R. Fergus · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Recurrent neural network regularization
W. Zaremba, I. Sutskever, and O. Vinyals · 2014
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Y. Gal and Z. Ghahramani · 2016
Cited alongside, same era.
Recurrent batch normalization
T. Cooijmans, N. Ballas, C. Laurent, and A. Courville · 2017
Closest in time.
Bayesian recurrent neural networks
M. Fortunato, C. Blundell, and O. Vinyals · 2017
Closest in time.
Improving neural language models with a continuous cache
E. Grave, A. Joulin, and N. Usunier · 2017
Closest in time.
Hypernetworks
D. Ha, A. Dai, and Q. Lee · 2017
Closest in time.
Tying word vectors and word classifiers: A loss framework for language modeling
H. Inan, K. Khosravi, and R. Socher · 2017
Closest in time.
Multiplicative LSTM for sequence modelling
B. Krause, I. Murray, S. Renals, and L. Lu · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Kalchbrenner, L. Espeholt, K. Simonyan, A. Oord, A. Graves, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Character-aware neural language models
Y. Kim, Y. Jernite, D. Sontag, and A. M. Rush · 2016
Cited alongside, same era.
Multiplicative LSTM for sequence modelling
B. Krause, L. Lu, I. Murray, and S. Renals · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
T. Salimans and D. P. Kingma · 2016
Cited alongside, same era.
On multiplicative integration with recurrent neural networks
Y. Wu, S. Zhang, Y. Zhang, Y. Bengio, and R. Salakhutdinov · 2016
Cited alongside, same era.
Gradual learning of deep recurrent neural networks
Z. Aharoni, G. Rattner, and H. Permuter · 2017
Cited alongside, same era.
Hierarchical multiscale recurrent neural networks
J. Chung, S. Ahn, and Y. Bengio · 2017
Cited alongside, same era.
G. Melis, C. Dyer, and P. Blunsom · 2017
Closest in time.
Fast-slow recurrent neural networks
A. Mujika, F. Meier, and A. Steger · 2017
Closest in time.
Learning simpler language models with the differential state framework
A. G. Ororbia II, T. Mikolov, and D. Reitter · 2017
Closest in time.
Using the output embedding to improve language models
O. Press and L. Wolf · 2017
Closest in time.
Recurrent highway networks
J. G. Zilly, R. K. Srivastava, J. Koutník, and J. Schmidhuber · 2017
Closest in time.
Neural architecture search with reinforcement learning
B. Zoph and Quoc V Le · 2017
Closest in time.