Fetching the paper…
Reading the bibliography…
Overparameterized deep networks have the capacity to memorize training data with zero \emph{training error}.
On the stability of inverse problems
Tikhonov, A. N · 1943
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Solutions of Ill Posed Problems
Tikhonov, A. N. and Arsenin, V. Y · 1977
Earlier work this paper cites.
Comparing biases for minimal network construction with back-propagation
Hanson, S. J. and Pratt, L. Y · 1988
Earlier work this paper cites.
Generalization and parameter estimation in feedforward nets: Some experiments
Morgan, N. and Bourlard, H · 1990
Earlier work this paper cites.
A simple weight decay can improve generalization
Krogh, A. and Hertz, J. A · 1992
Earlier work this paper cites.
Regularization and complexity control in feed-forward networks
Bishop, C. M · 1995
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Tibshirani, R · 1996
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Preventing “overfitting” of cross-validation data
Ng, A. Y · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Lecun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Overfitting in neural nets: Backpropagation, conjugate gradient, and early stopping
Caruana, R., Lawrence, S., and Giles, C. L · 2000
Earlier work this paper cites.
Statistical behavior and consistency of classification methods based on convex risk minimization
Zhang, T · 2004
Earlier work this paper cites.
80 million tiny images: A large data set for nonparametric object and scene recognition
Torralba, A., Fergus, R., and Freeman, W. T · 2008
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G · 2010
Earlier work this paper cites.
Pattern Recognition and Machine Learning (Information Science and Statistics)
Bishop, C. M · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y · 2011
Earlier work this paper cites.
Learning with noisy labels
Natarajan, N., Dhillon, I. S., Ravikumar, P. K., and Tewari, A · 2013
Earlier work this paper cites.
Dropout training as adaptive regularization
Wager, S., Wang, S., and Liang, P · 2013
Earlier work this paper cites.
Consistency of losses for learning from weak labels
Cid-Sueiro, J., García-García, D., and Santos-Rodríguez, R · 2014
Earlier work this paper cites.
Analysis of learning from positive and unlabeled data
du Plessis, M. C., Niu, G., and Sugiyama, M · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Cited alongside, same era.
Convex formulation for learning from positive and unlabeled data
du Plessis, M. C., Niu, G., and Sugiyama, M · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. L · 2015
Cited alongside, same era.
Introduction to statistical machine learning
Sugiyama, M · 2015
Cited alongside, same era.
A theory of learning with corrupted labels
van Rooyen, B. and Williamson, R. C · 2018
Later among the works it cites.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2018
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2019
Later among the works it cites.
MixMatch: A holistic approach to semi-supervised learning
Berthelot, D., Carlini, N., Goodfellow, I., Papernot, N., Oliver, A., and Raffel, C. A · 2019
Later among the works it cites.
Augmenting data with mixup for sentence classification: An empirical study
Guo, H., Mao, Y., and Zhang, R · 2019
Later among the works it cites.
Complementary-label learning for arbitrary losses and models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., and Shlens, J · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Arpit, D., Jastrzebski, S., Ballas, N., Krueger, D., Bengio, E., Kanwal, M. S., Maharaj, T., Fischer, A., Courville, A., Bengio, Y., and Lacoste-Julien, S · 2017
Cited alongside, same era.
Entropy-SGD: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2017
Cited alongside, same era.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Cited alongside, same era.
Ishida, T., Niu, G., Menon, A. K., and Sugiyama, M · 2019
Later among the works it cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Later among the works it cites.
SGD on neural networks learns functions of increasing complexity
Nakkiran, P., Kaplun, G., Kalimeris, D., Yang, T., Edelman, B. L., Zhang, F., and Barak, B · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
A meta-analysis of overfitting in machine learning
Roelofs, R., Shankar, V., Recht, B., Fridovich-Keil, S., Hardt, M., Miller, J., and Schmidt, L · 2019
Later among the works it cites.
A survey on image data augmentation for deep learning
Shorten, C. and Khoshgoftaar, T. M · 2019
Later among the works it cites.
On mixup training: Improved calibration and predictive uncertainty for deep neural networks
Thulasidasan, S., Chennupati, G., Bilmes, J., Bhattacharya, T., and Michalak, S · 2019
Later among the works it cites.
Interpolation consistency training for semi-supervised learning
Verma, V., Lamb, A., Kannala, J., Bengio, Y., and Lopez-Paz, D · 2019
Later among the works it cites.
Detecting overfitting via adversarial examples
Werpachowski, R., György, A., and Szepesvári, C · 2019
Later among the works it cites.
Sigua: Forgetting may make learning with noisy labels more robust
Han, B., Niu, G., Yu, X., Yao, Q., Xu, M., Tsang, I. W., and Sugiyama, M · 2020
Closest in time.
Proper ResNet implementation for CIFAR10/CIFAR100 in PyTorch
Idelbayev, Y · 2020
Closest in time.
Big transfer (BiT): General visual representation learning
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N · 2020
Closest in time.
Mitigating overfitting in supervised classification from two unlabeled datasets: A consistent risk correction approach
Lu, N., Zhang, T., Niu, G., and Sugiyama, M · 2020
Closest in time.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 2020
Closest in time.