Fetching the paper…
Reading the bibliography…
Convergence and convergence rate analyses of adaptive methods, such as Adaptive Moment Estimation (Adam) and its variants, have been widely studied for nonconvex optimization.
The Annals of Mathematical Statistics
Robbins, H., Monro, H.: A stochastic approximation method · 1951
Earlier work this paper cites.
USSR Computational Mathematics and Mathematical Physics
Polyak, B.T.: Some methods of speeding up the convergence of iteration methods · 1964
Earlier work this paper cites.
Doklady AN USSR
Nesterov, Y.: A method for unconstrained convex minimization problem with the rate of convergence · 1983
Earlier work this paper cites.
Cambridge University Press, Cambridge (1985)
Horn, R.A., Johnson, C.R.: Matrix Analysis · 1985
Earlier work this paper cites.
In: Proceedings of the 20th International Conference on Machine Learning, pp. 928–936 (2003)
Zinkevich, M.: Online convex programming and generalized infinitesimal gradient ascent · 2003
Earlier work this paper cites.
SIAM Journal on Optimization
Nemirovski, A., Juditsky, A., Lan, G., Shapiro, A.: Robust stochastic approximation approach to stochastic programming · 2009
Earlier work this paper cites.
In: Advances in Neural Information Processing Systems, vol. 23 (2010)
Zinkevich, M., Weimer, M., Li, L., Smola, A.: Parallelized stochastic gradient descent · 2010
Earlier work this paper cites.
Journal of Machine Learning Research
Duchi, J., Hazan, E., Singer, Y.: Adaptive subgradient methods for online learning and stochastic optimization · 2011
Earlier work this paper cites.
SIAM Journal on Optimization
Ghadimi, S., Lan, G.: Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization I: A generic algorithmic framework · 2012
Earlier work this paper cites.
COURSERA: Neural networks for machine learning
Tieleman, T., Hinton, G.: RMSProp: Divide the gradient by a running average of its recent magnitude · 2012
Earlier work this paper cites.
SIAM Journal on Optimization
Ghadimi, S., Lan, G.: Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization II: Shrinking procedures and optimal algorithms · 2013
Earlier work this paper cites.
In: Proceedings of The International Conference on Learning Representations (2015)
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization · 2015
Earlier work this paper cites.
In: Proceedings of the 32nd International Conference on Machine Learning, vol. 37, pp. 2048–2057 (2015)
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., Bengio, Y.: Show, attend and tell: Neural image caption generation with visual attention · 2015
Cited alongside, same era.
Arjovsky, M., Chintala, S., Bottou, L.: Wasserstein GAN
2017
Cited alongside, same era.
In: Advances in Neural Information Processing Systems, vol. 30 (2017)
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is All you Need · 2017
Cited alongside, same era.
SIAM Review
Bottou, L., Curtis, F.E., Nocedal, J.: Optimization methods for large-scale machine learning · 2018
Cited alongside, same era.
In: Proceedings of The International Conference on Learning Representations (2018)
Reddi, S.J., Kale, S., Kumar, S.: On the convergence of Adam and beyond · 2018
Cited alongside, same era.
In: Advances in Neural Information Processing Systems, vol. 32 (2019)
Zhang, G., Li, L., Nado, Z., Martens, J., Sachdeva, S., Dahl, G.E., Shallue, C.J., Grosse, R.: Which algorithmic choices matter at which batch sizes? Insights from a noisy quadratic model · 2019
Later among the works it cites.
In: Computer Vision and Pattern Recognition Conference, pp. 11,127–11,135 (2019)
Zou, F., Shen, L., Jie, Z., Zhang Weizhong, W.L.: A sufficient condition for convergences of Adam and RMSProp · 2019
Later among the works it cites.
In: Advances in Neural Information Processing Systems, vol. 33 (2020)
Chen, H., Zheng, L., AL Kontar, R., Raskutti, G.: Stochastic gradient descent in correlated settings: A study on Gaussian processes · 2020
Later among the works it cites.
Journal of Machine Learning Research
Fehrman, B., Gess, B., Jentzen, A.: Convergence rates for the stochastic gradient descent method for non-convex objective functions · 2020
Later among the works it cites.
In: Advances in Neural Information Processing Systems, vol. 33 (2020)
Mendler-Dünner, C., Perdomo, J.C., Zrnic, T., Hardt, M.: Stochastic optimization for performative prediction · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
In: Proceedings of The International Conference on Learning Representations (2018)
Smith, S.L., Kindermans, P.J., Le, Q.V.: Don’t decay the learning rate, increase the batch size · 2018
Cited alongside, same era.
In: Advances in Neural Information Processing Systems, vol. 31 (2018)
Virmaux, A., Scaman, K.: Lipschitz regularity of deep neural networks: analysis and efficient estimation · 2018
Cited alongside, same era.
In: S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, R. Garnett (eds.) Advances in Neural Information Processing Systems, vol. 31. Curran Associates, Inc. (2018)
Zaheer, M., Reddi, S., Sachan, D., Kale, S., Kumar, S.: Adaptive methods for nonconvex optimization · 2018
Cited alongside, same era.
In: Proceedings of The International Conference on Learning Representations (2019)
Chen, X., Liu, S., Sun, R., Hong, M.: On the convergence of a class of Adam-type algorithms for non-convex optimization · 2019
Cited alongside, same era.
In: International Conference on Learning Representations (2019)
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization · 2019
Cited alongside, same era.
In: Proceedings of The International Conference on Learning Representations (2019)
Luo, L., Xiong, Y., Liu, Y., Sun, X.: Adaptive gradient methods with dynamic bound of learning rate · 2019
Cited alongside, same era.
Journal of Machine Learning Research
Shallue, C.J., Lee, J., Antognini, J., Sohl-Dickstein, J., Frostig, R., Dahl, G.E.: Measuring the effects of data parallelism on neural network training · 2019
Cited alongside, same era.
Later among the works it cites.
In: Advances in Neural Information Processing Systems, vol. 33 (2020)
Scaman, K., Malherbe, C.: Robustness analysis of non-convex stochastic gradient descent using biased expectations · 2020
Later among the works it cites.
In: 12th Annual Workshop on Optimization for Machine Learning (2020)
Zhou, D., Chen, J., Cao, Y., Tang, Y., Yang, Z., Gu, Q.: On the convergence of adaptive gradient methods for nonconvex optimization · 2020
Later among the works it cites.
In: Advances in Neural Information Processing Systems, vol. 33 (2020)
Zhuang, J., Tang, T., Ding, Y., Tatikonda, S., Dvornek, N., Papademetris, X., Duncan, J.S.: AdaBelief optimizer: Adapting stepsizes by the belief in observed gradients · 2020
Later among the works it cites.
In: Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, vol. 452, pp. 3267–3275 (2021)
Chen, J., Zhou, D., Tang, Y., Yang, Z., Cao, Y., Gu, Q.: Closing the generalization gap of adaptive gradient methods in training deep neural network · 2021
Later among the works it cites.
In: Proceedings of the 24th International Conference on Artificial Intelligence and Statistics, vol. 130 (2021)
Gower, R.M., Sebbouh, O., Loizou, N.: SGD for structured nonconvex functions: Learning rates, minibatching and interpolation · 2021
Later among the works it cites.
IEEE Transactions on Cybernetics
Iiduka, H.: Appropriate learning rates of adaptive learning rate optimization algorithms for training deep neural networks · 2021
Later among the works it cites.
In: Proceedings of the 24th International Conference on Artificial Intelligence and Statistics, vol. 130 (2021)
Loizou, N., Vaswani, S., Laradji, I., Lacoste-Julien, S.: Stochastic polyak step-size for SGD: An adaptive learning rate for fast convergence · 2021
Later among the works it cites.