Fetching the paper…
Reading the bibliography…
Contrary to most machine learning models, modern deep artificial neural networks typically include multiple components that contribute to regularization.
On solving ill-posed problem and method of regularization
Tikhonov, A. N · 1963
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
Vapnik, V. N. and Chervonenkis, A. Y · 1971
Earlier work this paper cites.
Comparing biases for minimal network construction with back-propagation
Hanson, S. J. and Pratt, L. Y · 1989
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
LeCun, Y., Boser, B. E., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W. E., and Jackel, L. D · 1990
Earlier work this paper cites.
Bootstrap methods: another look at the jackknife
Efron, B · 1992
Earlier work this paper cites.
Training with noise is equivalent to Tikhonov regularization
Bishop, C. M · 1995
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2002
Earlier work this paper cites.
On early stopping in gradient descent learning
Yao, Y., Rosasco, L., and Caponnetto, A · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Deep sparse rectifier neural networks
Glorot, X., Bordes, A., and Bengio, Y · 2011
Earlier work this paper cites.
Unsupervised and transfer learning challenge: a deep learning approach
Mesnil, G. et al · 2011
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Maxout networks
Goodfellow, I. J., Warde-Farley, D., Mirza, M., Courville, A. C., and Bengio, Y · 2013
Earlier work this paper cites.
Dropout training as adaptive regularization
Wager, S., Wang, S., and Liang, P. S · 2013
Earlier work this paper cites.
Graham, B · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. R · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M. et al · 2015
Earlier work this paper cites.
Bouthillier, X., Konda, K., Vincent, P., and Memisevic, R · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
ImageNet large scale visual recognition challenge
Russakovsky, O. et al · 2015
Cited alongside, same era.
An analysis of deep neural network models for practical applications
Canziani, A., Paszke, A., and Culurciello, E · 2016
Cited alongside, same era.
Deep learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Do deep nets really need weight decay and dropout?
Hernández-García, A. and König, P · 2018
Closest in time.
Deep neural networks trained with heavier data augmentation learn features closer to representations in hIT
Hernández-García, A., Mehrer, J., Kriegeskorte, N., König, P., and Kietzmann, T. C · 2018
Closest in time.
Improving DNN robustness to adversarial attacks using Jacobian regularization
Jakubovitz, D. and Giryes, R · 2018
Closest in time.
Martin, C. H. and Mahoney, M. W · 2018
Closest in time.
Dropout training, data-dependent regularization, and generalization bounds
Mou, W., Zhou, Y., Gao, J., and Wang, L · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q · 2016
Cited alongside, same era.
FractalNet: Ultra-deep neural networks without residuals
Larsson, G., Maire, M., and Shakhnarovich, G · 2016
Cited alongside, same era.
Matching networks for one shot learning
Vinyals, O., Blundell, C., Lillicrap, T., Wierstra, D., et al · 2016
Cited alongside, same era.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M. J · 2017
Cited alongside, same era.
SGD learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A., Globerson, A., Malach, E., and Shalev-Shwartz, S · 2017
Cited alongside, same era.
Sensitivity and generalization in neural networks: an empirical study
Novak, R., Bahri, Y., Abolafia, D. A., Pennington, J., and Sohl-Dickstein, J · 2018
Closest in time.
How does batch normalization help optimization?
Santurkar, S., Tsipras, D., Ilyas, A., and Madry, A · 2018
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M., Javanmard, A., and Lee, J. D · 2018
Closest in time.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2019
Closest in time.
Invariance reduces variance: Understanding data augmentation in deep learning and beyond
Chen, S., Dobriban, E., and Lee, J. H · 2019
Closest in time.
Generalisation dynamics of online learning in over-parameterised neural networks
Goldt, S., Advani, M. S., Saxe, A. M., Krzakala, F., and Zdeborova, L · 2019
Closest in time.
Mixup as locally linear out-of-manifold regularization
Guo, H., Mao, Y., and Zhang, R · 2019
Closest in time.
Quantifying the carbon emissions of machine learning
Lacoste, A., Luccioni, A., Schmidt, V., and Dandres, T · 2019
Closest in time.
Does data augmentation lead to positive margin?
Rajput, S., Feng, Z., Charles, Z., Loh, P.-L., and Papailiopoulos, D · 2019
Closest in time.
A practical introduction to the bootstrap: a versatile method to make inferences by using data-driven simulations
Rousselet, G., Pernet, C., and Wilcox, R. R · 2019
Closest in time.
Energy and policy considerations for deep learning in nlp
Strubell, E., Ganesh, A., and McCallum, A · 2019
Closest in time.
EfficientNet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q. V · 2019
Closest in time.
Fixing the train-test resolution discrepancy
Touvron, H., Vedaldi, A., Douze, M., and Jégou, H · 2019
Closest in time.
Direct fit to nature: An evolutionary perspective on biological and artificial neural networks
Hasson, U., Nastase, S. A., and Goldstein, A · 2020
Closest in time.
Green algorithms: Quantifying the carbon emissions of computation
Lannelongue, L., Grealey, J., and Inouye, M · 2020
Closest in time.
Increasing the robustness of dnns against image corruptions by playing the game of noise
Rusak, E., Schott, L., Zimmermann, R., Bitterwolf, J., Bringmann, O., Bethge, M., and Brendel, W · 2020
Closest in time.