Fetching the paper…
Reading the bibliography…
Selecting an optimizer is a central step in the contemporary deep learning pipeline.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T · 1964
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate O(1/kˆ2)
Nesterov, Y. E · 1983
Earlier work this paper cites.
Improving the convergence of the backpropagation learning with second order methods
Becker, S. and Le Cun, Y · 1988
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Random search for hyper-parameter optimization
Bergstra, J. and Bengio, Y · 2012
Earlier work this paper cites.
Practical Bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Earlier work this paper cites.
Lecture 6.5-RMSProp: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., and Robinson, T · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Striving for simplicity: the all convolutional net
Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M · 2014
Earlier work this paper cites.
Batch normalization: accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: a method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
Incorporating Nesterov momentum into Adam
Dozat, T · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Cited alongside, same era.
Dissecting Adam: The sign, magnitude and variance of stochastic gradients
Balles, L. and Hennig, P · 2017
Cited alongside, same era.
Critical hyper-parameters: no random, no cry
Bousquet, O., Gelly, S., Kurach, K., Teytaud, O., and Vincent, D · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: training ImageNet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning. 2017
Lectures on convex optimization , volume 137
Nesterov, Y · 2018
Later among the works it cites.
Smith, L. N · 2018
Later among the works it cites.
Adaptive methods for nonconvex optimization
Zaheer, M., Reddi, S., Sachan, D., Kale, S., and Kumar, S · 2018
Later among the works it cites.
Adashift: Decorrelation and convergence of adaptive learning rate methods
Zhou, Z., Zhang, Q., Lu, G., Wang, H., Zhang, W., and Yu, Y · 2018
Later among the works it cites.
On the convergence of weighted AdaGrad with momentum for training deep neural networks
Zou, F. and Shen, L · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hessel, M., Modayil, J., and van Hasselt, H · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Cited alongside, same era.
Improving generalization performance by switching from Adam to SGD
Keskar, N. S. and Socher, R · 2017
Cited alongside, same era.
Fixing weight decay regularization in Adam
Loshchilov, I. and Hutter, F · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Cited alongside, same era.
Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Cited alongside, same era.
Yellowfin and the art of momentum tuning
Zhang, J. and Mitliagkas, I · 2017
Cited alongside, same era.
Adolphs, L., Kohler, J., and Lucchi, A · 2019
Closest in time.
NAMSG: An efficient method for training neural networks
Chen, J., Zhao, L., Qiao, X., and Fu, Y · 2019
Closest in time.
On the variance of the adaptive learning rate and beyond
Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J · 2019
Closest in time.
RoBERTa: a robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Closest in time.
Adaptive gradient methods with dynamic bound of learning rate
Luo, L., Xiong, Y., Liu, Y., and Sun, X · 2019
Closest in time.
On the convergence of Adam and beyond
Reddi, S. J., Kale, S., and Kumar, S · 2019
Closest in time.
Domain-independent dominance of adaptive methods
Savarese, P., McAllester, D., Babu, S., and Maire, M · 2019
Closest in time.
DeepOBS: a deep learning optimizer benchmark suite
Schneider, F., Balles, L., and Hennig, P · 2019
Closest in time.
Measuring the effects of data parallelism on neural network training
Shallue, C. J., Lee, J., Antognini, J., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2019
Closest in time.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q. V · 2019
Closest in time.
Mnasnet: Platform-aware neural architecture search for mobile
Tan, M., Chen, B., Pang, R., Vasudevan, V., Sandler, M., Howard, A., and Le, Q. V · 2019
Closest in time.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
Zhang, G., Li, L., Nado, Z., Martens, J., Sachdeva, S., Dahl, G. E., Shallue, C. J., and Grosse, R · 2019
Closest in time.
A sufficient condition for convergences of Adam and RMSProp
Zou, F., Shen, L., Jie, Z., Zhang, W., and Liu, W · 2019
Closest in time.