Fetching the paper…
Reading the bibliography…
Although adaptive optimization algorithms such as Adam show fast convergence in many machine learning tasks, this paper identifies a problem of Adam by analyzing its performance in a simple non-convex synthetic problem, showing that Adam's fast convergence would possibly lead the algorithm to local minimums.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B · 1964
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence o ( 1 / k 2 ) o(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database. in 2009 ieee conference on computer vision and pattern recognition
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Li, F.-F · 2009
Earlier work this paper cites.
Cifar-10 (canadian institute for advanced research)
Krizhevsky, A., Nair, V., and Hinton, G · 2009
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
Mcmahan, H. B. and Streeter, M · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Semantic contours from inverse detectors
Hariharan, B., Arbelaez, P., Bourdev, L., Maji, S., and Malik, J · 2011
Earlier work this paper cites.
Rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Adadelta: An adaptive learning rate method
Zeiler, M. D · 2012
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., and Robinson, T · 2013
Cited alongside, same era.
The pascal visual object classes challenge: A retrospective
Everingham, M., Eslami, S. M. A., Gool, L. V., Williams, C. K. I., Winn, J., and Zisserman, A · 2014
Cited alongside, same era.
Microsoft COCO: common objects in context
Lin, T., Maire, M., Belongie, S. J., Bourdev, L. D., Girshick, R. B., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. L · 2015
Cited alongside, same era.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
Chen, J. and Gu, Q · 2018
Later among the works it cites.
On the convergence of adam and beyond
Reddi, S. J., Kale, S., and Kumar., S · 2018
Later among the works it cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Later among the works it cites.
Adaptive methods for nonconvex optimization
Zaheer, M., Reddi, S., Sachan, D., Kale, S., and Kumar, S · 2018
Later among the works it cites.
On the convergence of adaptive gradient methods for nonconvex optimization
Zhou, D., Tang, Y., Yang, Z., Cao, Y., and Gu, Q · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs
Chen, L.-C., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A. L · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts, 2016
Loshchilov, I. and Hutter, F · 2016
Cited alongside, same era.
Accurate, large minibatch sgd: training imagenet in 1 hour
Goyal, P., Dollar, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Cited alongside, same era.
Pytorch gbw lm
Rdspring1
Cited in the paper.
Chen, X., Liu, S., Sun, R., and Hong, M · 2019
Later among the works it cites.
Nostalgic adam: Weighting more of the past gradients when designing the adaptive learning rate
Huang, H., Wang, C., and Dong., B · 2019
Later among the works it cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
Luo, L., Xiong, Y., Liu, Y., and Sun, X · 2019
Later among the works it cites.
Adashift: Decorrelation and convergence of adaptive learning rate methods
Zhou, Z., Zhang, Q., Lu, G., Wang, H., Zhang, W., and Yu, Y · 2019
Later among the works it cites.