Fetching the paper…
Reading the bibliography…
Dropout is a very effective way of regularizing neural networks.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. Denker, D. Henderson, R. Howard, W. Hubbard, and L. Jackel · 1989
Earlier work this paper cites.
Pruning from adaptive regularization
L. Hansen and C. Rasmussen · 1994
Earlier work this paper cites.
Training with noise is equivalent to Tikhonov regularization
C. M. Bishop · 1995
Earlier work this paper cites.
Regularization networks and support vector machines
T. Evgeniou, M. Pontil, and T. Poggio · 2000
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories
L. Fei-Fei, R. Fergus, and P. Perona · 2004
Earlier work this paper cites.
Caltech-256 object category dataset
G. Griffin, A. Holub, and P. Perona · 2007
Earlier work this paper cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Earlier work this paper cites.
Adaptive regularization of weight vectors
K. Crammer, A. Kulesza, and M. Dredze · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 2009
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng · 2011
Cited alongside, same era.
Adding noise to the input of a model trained with a regularized objective
S. Rifai, X. Glorot, B. Yoshua, and P. Vincent · 2011
Cited alongside, same era.
Unbiased look at dataset bias
A. Torralba and A. A. Efros · 2011
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
On fast dropout and its applicability to recurrent networks
J. Bayer, C. Osendorfer, and N. Chen · 2013
Annealed dropout training of deep networks
S. J. Rennie, V. Goel, and S. Thomas · 2014
Later among the works it cites.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Later among the works it cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Later among the works it cites.
Learning deep features for scene recognition using places database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Later among the works it cites.
Global optimality in tensor factorization, deep learning, and beyond
B. D. Haeffele and R. Vidal · 2015
Later among the works it cites.
Unsupervised Learning of Video Representations using LSTMs
N. Srivastava, E. Mansimov, and R. Salakhutdinov · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Robust subspace clustering
M. Soltanolkotabi, E. Elhamifar, and E. J. Candès · 2013
Cited alongside, same era.
Dropout training as adaptive regularization
S. Wager, S. Wang, and P. S. Liang · 2013
Cited alongside, same era.
Regularization of neural networks using dropconnect
L. Wan, M. Zeiler, S. Zhang, Y. L. Cun, and R. Fergus · 2013
Cited alongside, same era.
Fast dropout training
S. Wang and C. D. Manning · 2013
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Towards dropout training for convolutional neural networks
H. Wu and X. Gu
Cited in the paper.
Later among the works it cites.
Dropout training of matrix factorization and autoencoders for link prediction in sparse graphs
S. Zhai and Z. M. Zhang · 2015
Later among the works it cites.
Active Regression with Adaptive Huber loss
J. Cavazza and V. Murino · 2016
Later among the works it cites.
Adaptive dropout for training deep neural networks
B. Jimmy and B. Frey · 2016
Later among the works it cites.
Improved dropout for shallow and deep learning
Z. G. Li and T. Boqing Yang · 2016
Later among the works it cites.
A robust adaptive stochastic gradient method for deep learning
G. Caglar, S. Jose, M. Marcin, and Y. Bengio · 2017
Closest in time.