Fetching the paper…
Reading the bibliography…
Stochastic gradient descent with a large initial learning rate is widely used for training modern neural net architectures.
Sanjeev Arora, Simon S. Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
Can SGD learn recurrent neural networks with provable generalization?
Zeyuan Allen-Zhu and Yuanzhi Li · 1902
Earlier work this paper cites.
What can resnet learn efficiently, going beyond kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 1905
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop, coursera: Neural networks for machine learning
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
On the diffusion approximation of nonconvex stochastic gradient descent
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Later among the works it cites.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
Jinghui Chen and Quanquan Gu · 2018
Later among the works it cites.
Towards theoretical understanding of large batch training in stochastic gradient descent
Xiaowu Dai and Yuhua Zhu · 2018
Later among the works it cites.
Dnn’s sharpest directions along the sgd trajectory
Stanisław Jastrzębski, Zachary Kenton, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wenqing Hu, Chris Junchi Li, Lei Li, and Jian-Guo Liu · 2017
Cited alongside, same era.
Improving generalization performance by switching from adam to sgd
Nitish Shirish Keskar and Richard Socher · 2017
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix recovery
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2017
Cited alongside, same era.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L Smith and Quoc V Le · 2017
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
Large batch training of convolutional networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
Robert Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Later among the works it cites.
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Yoshua Bengio · 2018
Later among the works it cites.
The Step Decay Schedule: A Near Optimal, Geometrically Decaying Learning Rate Procedure
Rong Ge, Sham M. Kakade, Rahul Kidambi, and Praneeth Netrapalli · 2019
Closest in time.
Adaptive gradient methods with dynamic bound of learning rate
Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun · 2019
Closest in time.
Do deep neural networks learn shallow learnable examples first?
Karttikeya Mangalam and Vinay Prabhu · 2019
Closest in time.
SGD on Neural Networks Learns Functions of Increasing Complexity
Preetum Nakkiran, Gal Kaplun, Dimitris Kalimeris, Tristan Yang, Benjamin L. Edelman, Fred Zhang, and Boaz Barak · 2019
Closest in time.
Yeming Wen, Kevin Luk, Maxime Gazeau, Guodong Zhang, Harris Chan, and Jimmy Ba · 2019
Closest in time.