Fetching the paper…
Reading the bibliography…
We systematically explore regularizing neural networks by penalizing low entropy output distributions.
Information theory and statistical mechanics
Edwin T Jaynes · 1957
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
A maximum entropy approach to natural language processing
Adam L Berger, Vincent J Della Pietra, and Stephen A Della Pietra · 1996
Earlier work this paper cites.
A global optimization technique for statistical classifier design
David Miller, Ajit V Rao, Kenneth Rose, and Allen Gersho · 1996
Earlier work this paper cites.
Deterministic annealing for clustering, compression, classification, regression, and related optimization problems
Kenneth Rose · 1998
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely · 2011
Earlier work this paper cites.
Acoustic modeling using deep belief networks
Abdel-rahman Mohamed, George E Dahl, and Geoffrey Hinton · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann L Cun, and Rob Fergus · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
N-gram counts and language models from the common crawl
Christian Buck, Kenneth Heafield, and Bas Van Ooyen · 2014
Earlier work this paper cites.
Towards end-to-end speech recognition with recurrent neural networks
Alex Graves and Navdeep Jaitly · 2014
Cited alongside, same era.
Training deep neural networks on noisy labels with bootstrapping
Scott Reed, Honglak Lee, Dragomir Anguelov, Christian Szegedy, Dumitru Erhan, and Andrew Rabinovich · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Combining time-and frequency-domain convolution in convolutional neural network-based phone recognition
László Tóth · 2014
Cited alongside, same era.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Cited alongside, same era.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Martın Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al · 2016
Later among the works it cites.
End-to-end attention-based large vocabulary speech recognition
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio · 2016
Later among the works it cites.
Latent sequence decompositions
William Chan, Yu Zhang, Quoc Le, and Navdeep Jaitly · 2016
Later among the works it cites.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
William Chan, Navdeep Jaitly, Quoc V Le, and Oriol Vinyals · 2015
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
Sébastien Jean Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio · 2015
Cited alongside, same era.
Attention-based models for speech recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Gustav Larsson, Michael Maire, and Gregory Shakhnarovich · 2016
Later among the works it cites.
Learning online alignments with continuous rewards policy gradient
Yuping Luo, Chung-Cheng Chiu, Navdeep Jaitly, and Ilya Sutskever · 2016
Later among the works it cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Reward augmented maximum likelihood for neural structured prediction
Mohammad Norouzi, Dale Schuurmans, Samy Bengio, Zhifeng Chen, Navdeep Jaitly, Mike Schuster, and Yonghui Wu · 2016
Later among the works it cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2016
Later among the works it cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Later among the works it cites.
Disturblabel: Regularizing cnn on the loss layer
Lingxi Xie, Jingdong Wang, Zhen Wei, Meng Wang, and Qi Tian · 2016
Later among the works it cites.
Deep recurrent models with fast-forward connections for neural machine translation
Jie Zhou, Ying Cao, Xuguang Wang, Peng Li, and Wei Xu · 2016
Later among the works it cites.
Julian Georg Zilly, Rupesh Kumar Srivastava, Jan Koutník, and Jürgen Schmidhuber · 2016
Later among the works it cites.