Fetching the paper…
Reading the bibliography…
It is widely believed that a neural network can fit a training set containing at least as many samples as it has parameters, underpinning notions of overparameterized and underparameterized models.
Ridge regression: Biased estimation for nonorthogonal problems
Arthur E Hoerl and Robert W Kennard · 1970
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Principles of risk minimization for learning theory
Vladimir Vapnik · 1991
Earlier work this paper cites.
Neural net approximation
Andrew R Barron · 1992
Earlier work this paper cites.
Comparative accuracies of artificial neural networks and discriminant analysis in predicting forest cover types from cartographic variables
Jock A. Blackard and Denis J. Dean · 1999
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Li Deng · 2012
Earlier work this paper cites.
Approximation analysis of convolutional neural networks
Chenglong Bao, Qianxiao Li, Zuowei Shen, Cheng Tai, Lei Wu, and Xueshuang Xiang · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Deep vs. shallow networks: An approximation theory perspective
Hrushikesh N Mhaskar and Tomaso Poggio · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Cited alongside, same era.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Cited alongside, same era.
Provable approximation properties for deep neural networks
Uri Shaham, Alexander Cloninger, and Ronald R Coifman · 2018
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2018
Cited alongside, same era.
Shampoo: Preconditioned stochastic tensor optimization
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2021
Later among the works it cites.
Credit card dataset
Kaggle · 2021
Later among the works it cites.
Towards practical second order optimization for deep learning, 2021
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer · 2021
Later among the works it cites.
Convit: Improving vision transformers with soft convolutional inductive biases
Stéphane d’Ascoli, Hugo Touvron, Matthew L Leavitt, Ari S Morcos, Giulio Biroli, and Levent Sagun · 2021
Later among the works it cites.
Revisiting resnets: Improved training and scaling strategies
Irwan Bello, William Fedus, Xianzhi Du, Ekin Dogus Cubuk, Aravind Srinivas, Tsung-Yi Lin, Jonathon Shlens, and Barret Zoph · 2021
Later among the works it cites.
Stochastic training is not necessary for generalization
Jonas Geiping, Micah Goldblum, Phil Pope, Michael Moeller, and Tom Goldstein · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Cited alongside, same era.
Understanding generalization through visualizations
W Ronny Huang, Zeyad Emam, Micah Goldblum, Liam Fowl, JK Terry, Furong Huang, and Tom Goldstein · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Cited alongside, same era.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E Hinton · 2019
Cited alongside, same era.
Truth or backpropaganda? an empirical investigation of deep learning theory
Micah Goldblum, Jonas Geiping, Avi Schwarzschild, Michael Moeller, and Tom Goldstein · 2020
Cited alongside, same era.
Rethinking parameter counting in deep models: Effective dimensionality revisited
Wesley J Maddox, Gregory Benton, and Andrew Gordon Wilson · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Later among the works it cites.
Pac-bayes compression bounds so tight that they can explain generalization
Sanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski, Micah Goldblum, and Andrew G Wilson · 2022
Later among the works it cites.
Loss landscapes are all you need: Neural network generalization can be explained without the implicit bias of gradient descent
Ping-yeh Chiang, Renkun Ni, David Yu Miller, Arpit Bansal, Jonas Geiping, Micah Goldblum, and Tom Goldstein · 2022
Later among the works it cites.
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2022
Later among the works it cites.
Non-vacuous generalization bounds for large language models
Sanae Lotfi, Marc Finzi, Yilun Kuang, Tim GJ Rudner, Micah Goldblum, and Andrew Gordon Wilson · 2023
Later among the works it cites.
Efficient diffusion training via min-snr weighting strategy
Tiankai Hang, Shuyang Gu, Chen Li, Jianmin Bao, Dong Chen, Han Hu, Xin Geng, and Baining Guo · 2023
Later among the works it cites.
Efficiency 360: Efficient vision transformers
Badri N Patro and Vijay Agneeswaran · 2023
Later among the works it cites.
Comparing vision transformers and convolutional neural networks for image classification: A literature review
José Maurício, Inês Domingues, and Jorge Bernardino · 2023
Later among the works it cites.
Battle of the backbones: A large-scale comparison of pretrained models across computer vision tasks
Micah Goldblum, Hossein Souri, Renkun Ni, Manli Shu, Viraj Uday Prabhu, Gowthami Somepalli, Prithvijit Chattopadhyay, Mark Ibrahim, Adrien Bardes, Judy Hoffman, et al · 2023
Later among the works it cites.
Getting vit in shape: Scaling laws for compute-optimal model design
Ibrahim Alabdulmohsin, Xiaohua Zhai, Alexander Kolesnikov, and Lucas Beyer · 2023
Later among the works it cites.