Fetching the paper…
Reading the bibliography…
Regularization is one of the crucial ingredients of deep learning, yet the term regularization has various definitions, and regularization methods are often studied separately from each other.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition
Fukushima, K. and Miyake, S. (1982) · 1982
Earlier work this paper cites.
Experiments on learning by back propagation
Plaut, D. C., Nowlan, S. J., and Hinton, G. E. (1986) · 1986
Earlier work this paper cites.
Stochastic complexity and modeling
Rissanen, J. (1986) · 1986
Earlier work this paper cites.
Parallel distributed processing: Explorations in the microstructures of cognition. Volume 1: Foundations
Rumelhart, D. E., McClelland, J. L., and Group, P. R. (1986) · 1986
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W., and Jackel, L. D. (1989) · 1989
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Liu, D. C. and Nocedal, J. (1989) · 1989
Earlier work this paper cites.
Document image defect models
Baird, H. S. (1990) · 1990
Earlier work this paper cites.
Dimensionality reduction and prior knowledge in E-set recognition
Lang, K. J. and Hinton, G. E. (1990) · 1990
Earlier work this paper cites.
Generalization and parameter estimation in feedforward nets: Some experiments
Morgan, N. and Bourlard, H. (1990) · 1990
Earlier work this paper cites.
Generalization by weight-elimination with application to forecasting
Weigend, A. S., Rumelhart, D. E., and Huberman, B. A. (1991) · 1991
Earlier work this paper cites.
Simplifying neural networks by soft weight-sharing
Nowlan, S. J. and Hinton, G. E. (1992) · 1992
Earlier work this paper cites.
An efficient algorithm for learning invariance in adaptive classifiers
Simard, P., Le Cun, Y., Denker, J., and Victorri, B. (1992) · 1992
Earlier work this paper cites.
Network information criterion—determining the number of hidden units for an artificial neural network model
Murata, N., Yoshizawa, S., and Amari, S. (1994) · 1994
Earlier work this paper cites.
Regularization theory and neural networks architectures
Girosi, F., Jones, M., and Poggio, T. (1995) · 1995
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Hochreiter, S. and Schmidhuber, J. (1995) · 1995
Earlier work this paper cites.
The effects of adding noise during backpropagation training on a generalization performance
An, G. (1996) · 1996
Earlier work this paper cites.
Effective training of a neural network character classifier for word recognition
Yaegger, L., Lyon, R., and Webb, B. (1996) · 1996
Earlier work this paper cites.
Asymptotic statistical theory of overtraining and cross-validation
Amari, S., Murata, N., Muller, K.-R., Finke, M., and Yang, H. H. (1997) · 1997
Earlier work this paper cites.
Online algorithms and stochastic approximations
Bottou, L. (1998) · 1998
Earlier work this paper cites.
Multitask learning
Caruana, R. (1998) · 1998
Earlier work this paper cites.
Automatic early stopping using cross validation: quantifying the criteria
Prechelt, L. (1998) · 1998
Earlier work this paper cites.
Solving the ill-conditioning in neural network learning
van der Smagt, P. and Hirzinger, G. (1998) · 1998
Earlier work this paper cites.
Object recognition from local scale-invariant features
Lowe, D. G. (1999) · 1999
Earlier work this paper cites.
Training neural networks with additive noise in the desired signal
Wang, C. and Principe, J. C. (1999) · 1999
Earlier work this paper cites.
A model of inductive bias learning
Baxter, J. (2000) · 2000
Earlier work this paper cites.
Best practices for convolutional neural networks
Simard, P. Y., Steinkraus, D., and Platt, J. C. (2003) · 2003
Earlier work this paper cites.
Optimizing classifier performance via an approximation to the Wilcoxon-Mann-Whitney statistic
Yan, L., Dodier, R. H., Mozer, M., and Wolniewicz, R. H. (2003) · 2003
Earlier work this paper cites.
Model compression
Bucilă, C., Caruana, R., and Niculescu-Mizil, A. (2006) · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Hinton, G. E., Osindero, S., and Teh, Y.-W. (2006) · 2006
Earlier work this paper cites.
Principled hybrids of generative and discriminative models
Lasserre, J. A., Bishop, C. M., and Minka, T. P. (2006) · 2006
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Bengio, Y., Lamblin, P., Popovici, D., and Larochelle, H. (2007) · 2007
Earlier work this paper cites.
Training invariant support vector machines using selective sampling
Loosli, G., Canu, S., and Bottou, L. (2007) · 2007
Earlier work this paper cites.
Optimized approximation algorithm in neural networks without overfitting
Liu, Y., Starzyk, J. A., and Zhu, Z. (2008) · 2008
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J. (2009) · 2009
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Jarrett, K., Kavukcuoglu, K., LeCun, Y., et al. (2009) · 2009
Earlier work this paper cites.
Deep big simple neural nets excel on handwritten digit recognition
Ciresan, D. C., Meier, U., Gambardella, L. M., and Schmidhuber, J. (2010) · 2010
Cited alongside, same era.
Why does unsupervised pre-training help deep learning?
Erhan, D., Bengio, Y., Courville, A., Manzagol, P.-A., Vincent, P., and Bengio, S. (2010) · 2010
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y. (2010) · 2010
Cited alongside, same era.
Deep learning via Hessian-free optimization
Martens, J. (2010) · 2010
Cited alongside, same era.
Rectified linear units improve restricted Boltzmann machines
Nair, V. and Hinton, G. E. (2010) · 2010
Cited alongside, same era.
A survey on transfer learning
Pan, S. J. and Yang, Q. (2010) · 2010
Cited alongside, same era.
Semi-supervised learning with ladder networks
Rasmus, A., Berglund, M., Honkala, M., Valpola, H., and Raiko, T. (2015) · 2015
Later among the works it cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2015) · 2015
Later among the works it cites.
Gradual DropIn of layers to train very deep neural networks
Smith, L. N., Hand, E. M., and Doster, T. (2015) · 2015
Later among the works it cites.
Empirical evaluation of rectified activations in convolutional network
Xu, B., Wang, N., Chen, T., and Li, M. (2015) · 2015
Later among the works it cites.
Multi-scale context aggregation by dilated convolutions
Yu, F. and Koltun, V. (2015) · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep sparse rectifier neural networks
Glorot, X., Bordes, A., and Bengio, Y. (2011) · 2011
Cited alongside, same era.
On optimization methods for deep learning
Le, Q. V., Ngiam, J., Coates, A., Lahiri, A., Prochnow, B., and Ng, A. Y. (2011) · 2011
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2012) · 2012
Cited alongside, same era.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Cited alongside, same era.
Adaptive dropout for training deep neural networks
Ba, L. J. and Frey, B. (2013) · 2013
Cited alongside, same era.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P. (2013) · 2013
Cited alongside, same era.
Ba, J. L., Kiros, J. R., and Hinton, G. (2016) · 2016
Later among the works it cites.
Topology aware fully convolutional networks for histology gland segmentation
BenTaieb, A. and Hamarneh, G. (2016) · 2016
Later among the works it cites.
A guide to convolution arithmetic for deep learning
Dumoulin, V. and Visin, F. (2016) · 2016
Later among the works it cites.
Adaptive data augmentation for image classification
Fawzi, A., Horst, S., Turaga, D., and Frossard, P. (2016) · 2016
Later among the works it cites.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z. (2016) · 2016
Later among the works it cites.
Deep Learning
Goodfellow, I. J., Bengio, Y., and Courville, A. (2016) · 2016
Later among the works it cites.
Knowledge matters: Importance of prior information for optimization
Gülçehre, Ç. and Bengio, Y. (2016) · 2016
Later among the works it cites.
Train faster, generalize better: stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y. (2016) · 2016
Later among the works it cites.
Dreaming more data: Class-dependent distributions over diffeomorphisms for learned data augmentation
Hauberg, S., Freifeld, O., Larsen, A. B. L., Fisher III, J. W., and Hansen, L. K. (2016) · 2016
Later among the works it cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Later among the works it cites.
Generalizing and improving weight initialization
Hendrycks, D. and Gimpel, K. (2016) · 2016
Later among the works it cites.
Deep unsupervised learning through spatial contrasting
Hoffer, E., Hubara, I., and Ailon, N. (2016) · 2016
Later among the works it cites.
Perceptual losses for real-time style transfer and super-resolution
Johnson, J., Alahi, A., and Fei-Fei, L. (2016) · 2016
Later among the works it cites.
V-net: Fully convolutional neural networks for volumetric medical image segmentation
Milletari, F., Navab, N., and Ahmadi, S. A. (2016) · 2016
Later among the works it cites.
Regularization with stochastic transformations and perturbations for deep semi-supervised learning
Sajjadi, M., Javanmardi, M., and Tasdizen, T. (2016) · 2016
Later among the works it cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016) · 2016
Later among the works it cites.
Texture networks: Feed-forward synthesis of textures and stylized images
Ulyanov, D., Lebedev, V., Vedaldi, A., and Lempitsky, V. S. (2016) · 2016
Later among the works it cites.
Understanding data augmentation for classification: When to warp?
Wong, S. C., Gatt, A., Stamatescu, V., and McDonnell, M. D. (2016) · 2016
Later among the works it cites.
Synthesizing robust adversarial examples
Athalye, A. and Sutskever, I. (2017) · 2017
Closest in time.
Dataset augmentation in feature space
DeVries, T. and Taylor, G. W. (2017) · 2017
Closest in time.
Frerix, T., Möllenhoff, T., Moeller, M., and Cremers, D. (2017) · 2017
Closest in time.
Hoffer, E., Hubara, I., and Soudry, D. (2017) · 2017
Closest in time.
Deep nets don’t learn via memorization
Krueger, D., Ballas, N., Jastrzebski, S., Arpit, D., Kanwal, M. S., Maharaj, T., Bengio, E., Fischer, A., and Courville, A. (2017) · 2017
Closest in time.
No need to worry about adversarial examples in object detection in autonomous vehicles
Lu, J., Sibai, H., Fabry, E., and Forsyth, D. (2017) · 2017
Closest in time.
Morerio, P., Cavazza, J., Volpi, R., Vidal, R., and Murino, V. (2017) · 2017
Closest in time.
An overview of multi-task learning in deep neural networks
Ruder, S. (2017) · 2017
Closest in time.
Deep convolutional neural networks and data augmentation for environmental sound classification
Salamon, J. and Bello, J. P. (2017) · 2017
Closest in time.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B. (2017) · 2017
Closest in time.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2017) · 2017
Closest in time.
On over-fitting in model selection and subsequent selection bias in performance evaluation
Cawley, G. C. and Talbot, N. L. (2010) · 2079
Closest in time.