Fetching the paper…
Reading the bibliography…
Recurrent Neural Networks (RNNs) achieve state-of-the-art results in many sequence-to-sequence modeling tasks.
Long short-term memory
Sepp Hochreiter and J?rgen Schmidhuber · 1997
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
Hidetoshi Shimodaira · 2000
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Geoffrey E Hinton, Simon Osindero, and Yee-Whye Teh · 2006
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle · 2007
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2012
Earlier work this paper cites.
Training and analysing deep recurrent neural networks
Michiel Hermans and Benjamin Schrauwen · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
On the complexity of neural network classifiers: A comparison between shallow and deep architectures
Monica Bianchini and Franco Scarselli · 2014
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Guido Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Cited alongside, same era.
Net2net: Accelerating learning via knowledge transfer
Tianqi Chen, Ian J. Goodfellow, and Jonathon Shlens · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
David Ha, Andrew Dai, and Quoc V. Le · 2016
Later among the works it cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2016
Later among the works it cites.
Julian Georg Zilly, Rupesh Kumar Srivastava, Jan Koutn?k, and J?rgen Schmidhuber · 2016
Later among the works it cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V. Le · 2016
Later among the works it cites.
Dynamic evaluation of neural sequence models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Gradual dropin of layers to train very deep neural networks
Leslie N. Smith, Emily M. Hand, and Timothy Doster · 2015
Cited alongside, same era.
Tim Cooijmans, Nicolas Ballas, C?sar Laurent, ?aglar G?l?ehre, and Aaron Courville · 2016
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio
Cited in the paper.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio
Cited in the paper.
Closest in time.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom · 2017
Closest in time.
Regularizing and Optimizing LSTM Language Models
S. Merity, N. Shirish Keskar, and R. Socher · 2017
Closest in time.
Breaking the softmax bottleneck: a high-rank rnn language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W Cohen · 2017
Closest in time.