Fetching the paper…
Reading the bibliography…
Neural language models (LMs) based on recurrent neural networks (RNN) are some of the most successful word and character-level LMs.
Learning long-term dependencies with gradient descent is diffcult
Y Bengio, P Simard, and P Frasconi · 1994
Earlier work this paper cites.
Long short-term memory
S Hochreiter and J Schmidhuber · 1997
Earlier work this paper cites.
A neural probabilistic language model
Y Bengio, R Ducharme, P Vincent, and C Jauvin · 2003
Earlier work this paper cites.
Harnessing nonlinearity: Predicting chaotic systems and saving energy in wireless communication
H Jaeger and H Hass · 2004
Earlier work this paper cites.
Three new graphical models for statistical language modelling
A Mnih and G Hinton · 2007
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
M Guttmann and A Hyvarinen · 2010
Earlier work this paper cites.
Empirical evaluation and combination of advanced language modeling techniques
T Mikolov, A Deoras, S Kombrink, L Burget, and JH Cernocky · 2011
Cited alongside, same era.
Generating text with recurrent neural networks
I Sutskever, J Martens, and G Hinton · 2011
Cited alongside, same era.
The microsoft research sentence completion challenge
G Zweig and CJC Burges · 2011
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
GE Hinton, N Srivastava, A Krizhevsky, I Sutskever, and RR Salakhutdinov · 2012
Cited alongside, same era.
Subword language modelling with neural networks
T Mikolov, I Sutskever, A Deoras, HS Le, S Kombrink, and J Cernocky · 2012
Cited alongside, same era.
Understanding the exploding gradient problem
R Pascanu, T Mikolov, and Y Bengio · 2012
Later among the works it cites.
A neural autoregressive topic model
H Larochelle and S Lauly · 2012
Later among the works it cites.
A fast and simple algorithm for training neural probabilistic language models
A Mnih and YW Teh · 2012
Later among the works it cites.
Statistical language models based on neural networks
T Mikolov · 2012
Later among the works it cites.
On the importance of initialization and momentum in deep learning
I Sutskever, J Martens, G Dahl, and G Hinton · 2013
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…