Fetching the paper…
Reading the bibliography…
We show that dropout training is best understood as performing MAP estimation concurrently for a family of conditional models whose objectives are themselves lower bounded by the original dropout objective.
The psycho-biology of language
George Kingsley Zipf · 1935
Earlier work this paper cites.
Practical variational inference for neural networks
Alex Graves · 2011
Earlier work this paper cites.
Two problems with variational expectation maximisation for time-series models
Richard E Turner and Maneesh Sahani · 2011
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Earlier work this paper cites.
Understanding dropout
Pierre Baldi and Peter J Sadowski · 2013
Earlier work this paper cites.
On fast dropout and its applicability to recurrent networks
Justin Bayer, Christian Osendorfer, Daniela Korhammer, Nutan Chen, Sebastian Urban, and Patrick van der Smagt · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Regularization and nonlinearities for neural language models: when are they needed?
Marius Pachitariu and Maneesh Sahani · 2013
Earlier work this paper cites.
Fast dropout training
Sida Wang and Christopher Manning · 2013
Earlier work this paper cites.
An empirical analysis of dropout in piecewise linear networks
David Warde-Farley, Ian J Goodfellow, Aaron Courville, and Yoshua Bengio · 2013
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Cited alongside, same era.
Variational dropout and the local reparameterization trick
Diederik P Kingma, Tim Salimans, and Max Welling · 2015
Cited alongside, same era.
Dropout with expectation-linear regularization
Xuezhe Ma, Yingkai Gao, Zhiting Hu, Yaoliang Yu, Yuntian Deng, and Eduard H. Hovy · 2016
Cited alongside, same era.
Sharpening jensen’s inequality
JG Liao and Arthur Berg · 2017
Later among the works it cites.
Filtering variational objectives
Chris J Maddison, John Lawson, George Tucker, Nicolas Heess, Mohammad Norouzi, Andriy Mnih, Arnaud Doucet, and Yee Teh · 2017
Later among the works it cites.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom · 2017
Later among the works it cites.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2017
Later among the works it cites.
Regularizing neural networks by penalizing confident output distributions
Gabriel Pereyra, George Tucker, Jan Chorowski, Łukasz Kaiser, and Geoffrey Hinton · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Risk versus uncertainty in deep learning: Bayes, bootstrap and the dangers of dropout
Ian Osband · 2016
Cited alongside, same era.
Recurrent dropout without memory loss
Stanislau Semeniuta, Aliaksei Severyn, and Erhardt Barth · 2016
Cited alongside, same era.
Concrete dropout
Yarin Gal, Jiri Hron, and Alex Kendall · 2017
Cited alongside, same era.
Google vizier: A service for black-box optimization
Daniel Golovin, Benjamin Solnik, Subhodeep Moitra, Greg Kochanski, John Karro, and D Sculley · 2017
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani
Cited in the paper.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani
Cited in the paper.
Data noising as smoothing in neural network language models
Ziang Xie, Sida I Wang, Jiwei Li, Daniel Lévy, Aiming Nie, Dan Jurafsky, and Andrew Y Ng · 2017
Later among the works it cites.
Breaking the softmax bottleneck: a high-rank rnn language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W Cohen · 2017
Later among the works it cites.
Konrad Zolna, Devansh Arpit, Dendi Suhubdy, and Yoshua Bengio · 2017
Later among the works it cites.