Fetching the paper…
Reading the bibliography…
In natural language processing, a lot of the tasks are successfully solved with recurrent neural networks, but such models have a huge number of parameters.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. 1993 · 1993
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Extensions of recurrent neural network language model
T. Mikolov, S. Kombrink, L. Burget, J. Cernocky, and S. Khudanpur. 2011 · 2011
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Diederik P Kingma, Tim Salimans, and Max Welling. 2015 · 2015
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Quoc V. Le, Navdeep Jaitly, and Geoffrey E. Hinton. 2015 · 2015
Cited alongside, same era.
Exploring models and data for image question answering
Mengye Ren, Ryan Kiros, and Richard S. Zemel. 2015 · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
X. Zhang, J. Zhao, and Y. LeCun. 2015 · 2015
Cited alongside, same era.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals. 2016 · 2016
Cited alongside, same era.
Bayesian recurrent neural networks
Meire Fortunato, Charles Blundell, and Oriol Vinyals. 2017 · 2017
Later among the works it cites.
Hypernetworks
David Ha, Andrew Dai, and Quoc V. Le. 2017 · 2017
Later among the works it cites.
Bayesian compression for deep learning
Christos Louizos, Karen Ullrich, and Max Welling. 2017 · 2017
Later among the works it cites.
Variational dropout sparsifies deep neural networks
Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov. 2017 · 2017
Later among the works it cites.
Exploring sparsity in recurrent neural networks
Sharan Narang, Gregory F. Diamos, Shubho Sengupta, and Erich Elsen. 2017 · 2017
Later among the works it cites.
Structured bayesian pruning via log-normal multiplicative noise
Kirill Neklyudov, Dmitry Molchanov, Arsenii Ashukha, and Dmitry P Vetrov. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Cited alongside, same era.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Compressing recurrent neural network with tensor train
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura. 2017 · 2017
Later among the works it cites.
Learning intrinsic sparse structures within long short-term memory
Wei Wen, Yuxiong He, Samyam Rajbhandari, Minjia Zhang, Wenhan Wang, Fang Liu, Bin Hu, Yiran Chen, and Hai Li. 2018 · 2018
Closest in time.