Fetching the paper…
Reading the bibliography…
Dropout is a simple but effective technique for learning in neural networks and other settings.
Distribution inequalities for the binomial law
E. Slud · 1977
Earlier work this paper cites.
Concrete Mathematics
R. L. Graham, D. E. Knuth, and O. Patashnik · 1989
Earlier work this paper cites.
Stochastic approximation algorithms and applications
H. J. Kushner and G. G. Yin · 1997
Earlier work this paper cites.
A second-order perceptron algorithm
N. Cesa-Bianchi, A. Conconi, and C. Gentile · 2002
Earlier work this paper cites.
Some infinity theory for predictor ensembles
L. Breiman · 2004
Earlier work this paper cites.
Statistical behavior and consistency of classification methods based on convex risk minimization
T. Zhang · 2004
Earlier work this paper cites.
Convexity, classification, and risk bounds
P. L. Bartlett, M. I. Jordan, and J. D. McAuliffe · 2006
Earlier work this paper cites.
Asymptotic theory of statistics and probability
A. DasGupta · 2008
Earlier work this paper cites.
Random classification noise defeats all convex potential boosters
P. M. Long and R. A. Servedio · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Deep learning how i did it: Merck 1st place interview, 2012
G. Dahl · 2012
Cited alongside, same era.
On the necessity of irrelevant variables
D. P. Helmbold and P. M. Long · 2012
Cited alongside, same era.
Dropout: a simple and effective way to improve neural networks, 2012
G. E. Hinton · 2012
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors, 2012
G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov · 2012
Recent advances in deep learning for speech research at microsoft
e. a. L. Deng · 2013
Later among the works it cites.
Dropout training as adaptive regularization
S. Wager, S. I. Wang, and P. Liang · 2013
Later among the works it cites.
Regularization of neural networks using dropconnect
L. Wan, M. Zeiler, S. Zhang, Y. L. Cun, and R. Fergus · 2013
Later among the works it cites.
Fast dropout training
S. Wang and C. Manning · 2013
Later among the works it cites.
Online linear optimization via smoothing
J. Abernethy, C. Lee, A. Sinha, and A. Tewari · 2014
Closest in time.
Learning with pseudo-ensembles
P. Bachman, O. Alsharif, and D. Precup · 2014
Closest in time.
Follow the leader with dropout perturbations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Understanding dropout
P. Baldi and P. J. Sadowski · 2013
Cited alongside, same era.
Improving deep neural networks for lvcsr using rectified linear units and dropout
G. E. Dahl, T. N. Sainath, and G. E. Hinton · 2013
Cited alongside, same era.
T. Van Erven, W. Kotłowski, and M. K. Warmuth · 2014
Closest in time.
Altitude training: Strong bounds for single-layer dropout
S. Wager, W. Fithian, S. Wang, and P. S. Liang · 2014
Closest in time.