Fetching the paper…
Reading the bibliography…
In this paper, we describe a phenomenon, which we named "super-convergence", where neural networks can be trained an order of magnitude faster than with standard training methods.
A method of solving a convex programming problem with convergence rate o (1/k2)
Yurii Nesterov · 1983
Earlier work this paper cites.
Simulated annealing and boltzmann machines
Emile Aarts and Jan Korst · 1988
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Stochastic gradient descent tricks
Léon Bottou · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George E Dahl, and Geoffrey E Hinton · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann L Cun, and Rob Fergus · 2013
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Ian J Goodfellow, Oriol Vinyals, and Andrew M Saxe · 2014
Earlier work this paper cites.
Adasecant: robust adaptive secant method for stochastic gradient
Caglar Gulcehre, Marcin Moczulski, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
No more pesky learning rate guessing games
Leslie N. Smith · 2015
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, and Yann LeCun · 2016
Pratik Chaudhari and Stefano Soatto · 2017
Closest in time.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, et al · 2017
Closest in time.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Closest in time.
A robust adaptive stochastic gradient method for deep learning
Caglar Gulcehre, Jose Sotelo, Marcin Moczulski, and Yoshua Bengio · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
An empirical analysis of deep network loss surfaces
Daniel Jiwoong Im, Michael Tao, and Kristin Branson · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Qianli Liao, Kenji Kawaguchi, and Tomaso Poggio · 2016
Cited alongside, same era.
Sgdr: stochastic gradient descent with restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Closest in time.
Three factors influencing minima in sgd
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Closest in time.
Generalization in deep learning
Kenji Kawaguchi, Leslie Pack Kaelbling, and Yoshua Bengio · 2017
Closest in time.
Cyclical learning rates for training neural networks
Leslie N. Smith · 2017
Closest in time.
Understanding generalization and stochastic gradient descent
Samuel L Smith and Quoc V Le · 2017
Closest in time.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, and Quoc V Le · 2017
Closest in time.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi · 2017
Closest in time.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, et al · 2017
Closest in time.
Leslie N. Smith · 2018
Closest in time.