Fetching the paper…
Reading the bibliography…
The noise in stochastic gradient descent (SGD) provides a crucial implicit regularization effect for training overparameterized models.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
On generalization error bounds of noisy gradient methods for non-convex learning
Jian Li, Xuanyuan Luo, and Mingda Qiao · 1902
Earlier work this paper cites.
Exponential convergence of langevin distributions and their discrete approximations
Gareth O Roberts, Richard L Tweedie, et al · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
Concentration inequalities and martingale inequalities: a survey
Fan Chung and Linyuan Lu · 2006
Earlier work this paper cites.
Mcmc using hamiltonian dynamics
Radford M Neal et al · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Minimax-optimal rates for sparse additive models over kernel classes via convex programming
Garvesh Raskutti, Martin J Wainwright, and Bin Yu · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Living on the edge: Phase transitions in convex programs with random data
Dennis Amelunxen, Martin Lotz, Michael B McCoy, and Joel A Tropp · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adding gradient noise improves learning for very deep networks
Arvind Neelakantan, Luke Vilnis, Quoc V Le, Ilya Sutskever, Lukasz Kaiser, Karol Kurach, and James Martens · 2015
Earlier work this paper cites.
Path-sgd: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Russ R Salakhutdinov, and Nati Srebro · 2015
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Consistency and fluctuations for stochastic gradient langevin dynamics
Yee Whye Teh, Alexandre H Thiery, and Sebastian J Vollmer · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Theoretical guarantees for approximate sampling from smooth and log-concave densities
Arnak S Dalalyan · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2017
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2019
Later among the works it cites.
Learning imbalanced datasets with label-distribution-aware margin loss
Kaidi Cao, Colin Wei, Adrien Gaidon, Nikos Arechiga, and Tengyu Ma · 2019
Later among the works it cites.
Nonconvex optimization meets low-rank matrix factorization: An overview
Yuejie Chi, Yue M Lu, and Yuxin Chen · 2019
Later among the works it cites.
Limitations of lazy training of two-layers neural network
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Later among the works it cites.
The implicit bias of depth: How incremental learning drives generalization
Daniel Gissin, Shai Shalev-Shwartz, and Amit Daniely · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generalization bounds of sgld for non-convex learning: Two theoretical viewpoints
Wenlong Mou, Liwei Wang, Xiyu Zhai, and Kai Zheng · 2017
Cited alongside, same era.
Theory of deep learning iii: explaining the non-overfitting puzzle
Tomaso Poggio, Kenji Kawaguchi, Qianli Liao, Brando Miranda, Lorenzo Rosasco, Xavier Boix, Jack Hidary, and Hrushikesh Mhaskar · 2017
Cited alongside, same era.
Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis
Maxim Raginsky, Alexander Rakhlin, and Matus Telgarsky · 2017
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
A hitting time analysis of stochastic gradient langevin dynamics
Yuchen Zhang, Percy Liang, and Moses Charikar · 2017
Cited alongside, same era.
Theoretical analysis of auto rate-tuning by batch normalization
Sanjeev Arora, Zhiyuan Li, and Kaifeng Lyu · 2018
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Later among the works it cites.
On dropout and nuclear norm regularization
Poorya Mianjy and Raman Arora · 2019
Later among the works it cites.
Lexicographic and depth-sensitive margins in homogeneous and non-homogeneous deep models
Mor Shpigel Nacson, Suriya Gunasekar, Jason D Lee, Nathan Srebro, and Daniel Soudry · 2019
Later among the works it cites.
Information-theoretic generalization bounds for sgld via data-dependent estimates
Jeffrey Negrea, Mahdi Haghifam, Gintare Karolina Dziugaite, Ashish Khisti, and Daniel M Roy · 2019
Later among the works it cites.
Implicit regularization for optimal sparse recovery
Tomas Vaskevicius, Varun Kanade, and Patrick Rebeschini · 2019
Later among the works it cites.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
Sharan Vaswani, Francis Bach, and Mark Schmidt · 2019
Later among the works it cites.
Data-dependent sample complexity of deep neural networks via lipschitz augmentation
Colin Wei and Tengyu Ma · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2019
Later among the works it cites.
How noise affects the hessian spectrum in overparameterized neural networks
Mingwei Wei and David J Schwab · 2019
Later among the works it cites.
Yeming Wen, Kevin Luk, Maxime Gazeau, Guodong Zhang, Harris Chan, and Jimmy Ba · 2019
Later among the works it cites.
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Later among the works it cites.
Disentangling adaptive gradient methods from learning rates
Naman Agarwal, Rohan Anil, Elad Hazan, Tomer Koren, and Cyril Zhang · 2020
Closest in time.
Dropout: Explicit forms and capacity control
Raman Arora, Peter Bartlett, Poorya Mianjy, and Nathan Srebro · 2020
Closest in time.
Optimal regularization can mitigate double descent
Preetum Nakkiran, Prayaag Venkat, Sham Kakade, and Tengyu Ma · 2020
Closest in time.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Closest in time.
Implicit regularization in deep learning may not be explainable by norms
Noam Razin and Nadav Cohen · 2020
Closest in time.
The implicit and explicit regularization effects of dropout
Colin Wei, Sham Kakade, and Tengyu Ma · 2020
Closest in time.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Closest in time.