Fetching the paper…
Reading the bibliography…
Ongoing innovations in recurrent neural network architectures have provided a steady influx of apparently state-of-the-art results on language modelling benchmarks.
Building a large annotated corpus of english: The Penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Violin plots: a box plot-density trace synergism
Jerry L Hintze and Ray D Nelson · 1998
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
The human knowledge compression contest
Marcus Hutter · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
Parallelizing exploration-exploitation tradeoffs in Gaussian process bandit optimization
Thomas Desautels, Andreas Krause, and Joel W. Burdick · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Hasim Sak, Andrew W. Senior, and Françoise Beaufays · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Earlier work this paper cites.
Nal Kalchbrenner, Ivo Danihelka, and Alex Graves · 2015
Cited alongside, same era.
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio · 2016
Cited alongside, same era.
Capacity and trainability in recurrent neural networks
Jasmine Collins, Jascha Sohl-Dickstein, and David Sussillo · 2016
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier · 2016
Recurrent dropout without memory loss
Stanislau Semeniuta, Aliaksei Severyn, and Erhardt Barth · 2016
Later among the works it cites.
On multiplicative integration with recurrent neural networks
Yuhuai Wu, Saizheng Zhang, Ying Zhang, Yoshua Bengio, and Ruslan Salakhutdinov · 2016
Later among the works it cites.
Julian G. Zilly, Rupesh Kumar Srivastava, Jan Koutník, and Jürgen Schmidhuber · 2016
Later among the works it cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Later among the works it cites.
Google vizier: A service for black-box optimization
Daniel Golovin, Benjamin Solnik, Subhodeep Moitra, Greg Kochanski, John Karro, and D Sculley · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher · 2016
Cited alongside, same era.
Neural machine translation in linear time
Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aäron van den Oord, Alex Graves, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2016
Cited alongside, same era.
Closest in time.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2017
Closest in time.
Dynamic evaluation of neural sequence models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals · 2017
Closest in time.
Regularizing and optimizing LSTM language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2017
Closest in time.
Nils Reimers and Iryna Gurevych · 2017
Closest in time.