Fetching the paper…
Reading the bibliography…
The recently introduced dropout training criterion for neural networks has been the subject of much attention due to its simplicity and remarkable effectiveness as a regularizer, as well as its interpretation as a training procedure for an exponentially large ensemble of networks that share parameters.
A refinement of the arithmetic mean-geometric mean inequality
Cartwright, D. I. and Field, M. J. (1978) · 1978
Earlier work this paper cites.
The strength of weak learnability
Schapire, R. E. (1990) · 1990
Earlier work this paper cites.
Bagging predictors
Breiman, L. (1994) · 1994
Earlier work this paper cites.
Training with noise is equivalent to Tikhonov regularization
Bishop, C. M. (1995) · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998) · 1998
Earlier work this paper cites.
Popular ensemble methods: An empirical study
Opitz, D. and Maclin, R. (1999) · 1999
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Jarrett, K., Kavukcuoglu, K., Ranzato, M., and LeCun, Y. (2009) · 2009
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
Bergstra, J., Breuleux, O., Bastien, F., Lamblin, P., Pascanu, R., Desjardins, G., Turian, J., Warde-Farley, D., and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Vincent, P., Larochelle, H., Lajoie, I., Bengio, Y., and Manzagol, P.-A. (2010) · 2010
Cited alongside, same era.
Deep sparse rectifier neural networks
Glorot, X., Bordes, A., and Bengio, Y. (2011) · 2011
Cited alongside, same era.
The manifold tangent classifier
Rifai, S., Dauphin, Y., Vincent, P., Bengio, Y., and Muller, X. (2011) · 2011
Cited alongside, same era.
Theano: new features and speed improvements
Bastien, F., Lamblin, P., Pascanu, R., Bergstra, J., Goodfellow, I. J., Bergeron, A., Bouchard, N., and Bengio, Y. (2012) · 2012
Cited alongside, same era.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y. (2012) · 2012
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
Understanding dropout
Baldi, P. and Sadowski, P. J. (2013) · 2013
Closest in time.
Maxout networks
Goodfellow, I. J., Warde-Farley, D., Mirza, M., Courville, A., and Bengio, Y. (2013a) · 2013
Closest in time.
Improving Neural Networks With Dropout
Srivastava, N. (2013) · 2013
Closest in time.
Dropout training as adaptive regularization
Wager, S., Wang, S., and Liang, P. (2013) · 2013
Closest in time.
Regularization of neural networks using dropconnect
Wan, L., Zeiler, M., Zhang, S., LeCun, Y., and Fergus, R. (2013) · 2013
Closest in time.
Fast dropout training
Wang, S. and Manning, C. (2013) · 2013
Closest in time.
Stochastic pooling for regularization of deep convolutional neural networks
Zeiler, M. D. and Fergus, R. (2013) · 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinv, R. (2012) · 2012
Cited alongside, same era.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. (2012a) · 2012
Cited alongside, same era.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. (2012b) · 2012
Cited alongside, same era.
Pylearn2: a machine learning research library
Goodfellow, I. J., Warde-Farley, D., Lamblin, P., Dumoulin, V., Mirza, M., Pascanu, R., Bergstra, J., Bastien, F., and Bengio, Y. (2013b)
Cited in the paper.
Closest in time.
On rectified linear units for speech processing
Zeiler, M. D., Ranzato, M., Monga, R., Mao, M., Yang, K., Le, Q., Nguyen, P., Senior, A., Vanhoucke, V., Dean, J., and Hinton, G. E. (2013) · 2013
Closest in time.