Fetching the paper…
Reading the bibliography…
Training neural networks on image datasets generally require extensive experimentation to find the optimal learning rate regime.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 1910
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
D Maclaurin, D Duvenaud, and R.P Adams · 2015
Earlier work this paper cites.
Lets keep it simple, using simple architectures to outperform deeper and more complex architectures
Seyyed Hossein HasanPour, Mohammad Rouhani, Mohsen Fayyaz, and Mohammad Sabokrou · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Improved regularization of convolutional neural networks with cutout
Terrence DeVries and Graham W. Taylor · 2017
Cited alongside, same era.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Later among the works it cites.
Exploring loss function topology with cyclical learning rates
Leslie N Smith and Nicholay Topin · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
A.C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht · 2017
Later among the works it cites.
Online learning rate adaptation with hypergradient descent
Atilim Gunes Baydin, Robert Cornish, David Martinez Rubio, Mark Schmidt, and Frank Wood · 2018
Later among the works it cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Densely connected convolutional networks
Guo Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger · 2017
Cited alongside, same era.
Hyperband: Bandit-based configuration evaluation for hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar · 2017
Cited alongside, same era.
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2018
Later among the works it cites.
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Yoshua Bengio · 2018
Later among the works it cites.