Fetching the paper…
Reading the bibliography…
Intriguing empirical evidence exists that deep learning can work well with exoticschedules for varying the learning rate.
An elementary proof of a theorem of johnson and lindenstrauss
Sanjoy Dasgupta and Anupam Gupta · 2003
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
SGDR: Stochastic Gradient Descent with Warm Restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Instance normalization: The missing ingredient for fast stylization
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2016
Earlier work this paper cites.
Riemannian approach to batch normalization
Minhyung Cho and Jaehyung Lee · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Cited alongside, same era.
L2 regularization versus batch and weight normalization
Twan van Laarhoven · 2017
Cited alongside, same era.
Large Batch Training of Convolutional Networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
How does batch normalization help optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Later among the works it cites.
WNGrad: Learn the Learning Rate in Gradient Descent
Xiaoxia Wu, Rachel Ward, and Léon Bottou · 2018
Later among the works it cites.
Group normalization
Yuxin Wu and Kaiming He · 2018
Later among the works it cites.
Theoretical analysis of auto rate-tuning by batch normalization
Sanjeev Arora, Zhiyuan Li, and Kaifeng Lyu · 2019
Closest in time.
Sgd: General analysis and improved rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Closest in time.
Three mechanisms of weight decay regularization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nils Bjorck, Carla P Gomes, Bart Selman, and Kilian Q Weinberger · 2018
Cited alongside, same era.
Jonas Kohler, Hadi Daneshmand, Aurelien Lucchi, Ming Zhou, Klaus Neymeyr, and Thomas Hofmann · 2018
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun
Cited in the paper.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun
Cited in the paper.
Norm matters: efficient and accurate normalization schemes in deep networks
Elad Hoffer, Ron Banner, Itay Golan, and Daniel Soudry
Cited in the paper.
Fix your classifier: the marginal value of training the last weight layer
Elad Hoffer, Itay Hubara, and Daniel Soudry
Cited in the paper.
How to train your resnet 6: Weight decay?
David Page
Cited in the paper.
Guodong Zhang, Chaoqi Wang, Bowen Xu, and Roger Grosse · 2019
Closest in time.