Fetching the paper…
Reading the bibliography…
There is a notable dearth of results characterizing the preconditioning effect of Adam and showing how it may alleviate the curse of ill-conditioning -- an issue plaguing gradient descent (GD).
Applied numerical linear algebra
Demmel, J. W. (1997) · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I. (1998) · 1998
Earlier work this paper cites.
Fast and near-optimal diagonal preconditioning
Jambulapati, A., Li, J., Musco, C., Sidford, A., and Tian, K. (2020) · 2008
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K. (2012) · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Iteration complexity of randomized block-coordinate descent methods for minimizing a composite function
Richtárik, P. and Takáč, M. (2014) · 2014
Earlier work this paper cites.
Preconditioning
Wathen, A. J. (2015) · 2015
Earlier work this paper cites.
Coordinate descent algorithms
Wright, S. J. (2015) · 2015
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M. (2016) · 2016
Earlier work this paper cites.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Balles, L. and Hennig, P. (2018) · 2018
Earlier work this paper cites.
signsgd: Compressed optimisation for non-convex problems
Bernstein, J., Wang, Y.-X., Azizzadenesheli, K., and Anandkumar, A. (2018) · 2018
Cited alongside, same era.
On the convergence of a class of adam-type algorithms for non-convex optimization
Chen, X., Liu, S., Sun, R., and Hong, M. (2018) · 2018
Cited alongside, same era.
On the convergence of adam and beyond
Reddi, S. J., Kale, S., and Kumar, S. (2018) · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
Zaheer, M., Reddi, S., Sachan, D., Kale, S., and Kumar, S. (2018) · 2018
Cited alongside, same era.
On the convergence of adaptive gradient methods for nonconvex optimization
Zhou, D., Chen, J., Cao, Y., Tang, Y., Yang, Z., and Gu, Q. (2018) · 2018
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Super-adam: faster and universal framework of adaptive gradients
Huang, F., Li, J., and Huang, H. (2021) · 2021
Later among the works it cites.
Rmsprop converges with proper hyperparameter
Shi, N. and Li, D. (2021) · 2021
Later among the works it cites.
Adagrad avoids saddle points
Antonakopoulos, K., Mertikopoulos, P., Piliouras, G., and Wang, X. (2022) · 2022
Later among the works it cites.
Robustness to unbounded smoothness of generalized signsgd
Crawshaw, M., Liu, M., Orabona, F., Zhang, W., and Zhuang, Z. (2022) · 2022
Later among the works it cites.
A simple convergence proof of adam and adagrad
Défossez, A., Bottou, L., Bach, F., and Usunier, N. (2022) · 2022
Later among the works it cites.
The power of adaptivity in sgd: Self-tuning step sizes with unbounded gradients and affine variance
Faw, M., Tziotis, I., Caramanis, C., Mokhtari, A., Shakkottai, S., and Ward, R. (2022) · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J. (2019) · 2019
Cited alongside, same era.
On the convergence of stochastic gradient descent with adaptive stepsizes
Li, X. and Orabona, F. (2019) · 2019
Cited alongside, same era.
Escaping saddle points with adaptive gradient methods
Staib, M., Reddi, S., Kale, S., Kumar, S., and Sra, S. (2019) · 2019
Cited alongside, same era.
A general system of differential equations to model first-order adaptive algorithms
Da Silva, A. B. and Gazeau, M. (2020) · 2020
Cited alongside, same era.
New insights and perspectives on the natural gradient method
Martens, J. (2020) · 2020
Cited alongside, same era.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes
Ward, R., Wu, X., and Bottou, L. (2020) · 2020
Cited alongside, same era.
Later among the works it cites.
Wang, B., Zhang, Y., Zhang, H., Meng, Q., Ma, Z.-M., Liu, T.-Y., and Chen, W. (2022) · 2022
Later among the works it cites.
Adam can converge without any modification on update rules
Zhang, Y., Chen, C., Shi, N., Sun, R., and Luo, Z.-Q. (2022) · 2022
Later among the works it cites.
Beyond uniform smoothness: A stopped analysis of adaptive sgd
Faw, M., Rout, L., Caramanis, C., and Shakkottai, S. (2023) · 2023
Later among the works it cites.
Toward understanding why adam converges faster than sgd for transformers
Pan, Y. and Li, Y. (2023) · 2023
Later among the works it cites.
Convergence of adagrad for non-convex objectives: Simple proofs and relaxed assumptions
Wang, B., Zhang, H., Ma, Z., and Chen, W. (2023) · 2023
Later among the works it cites.