Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
H. Karimi, J. Nutini, and M. Schmidt · 2016
Cited alongside, same era.
Natasha 2: Faster non-convex optimization than sgd
Original
Z. Allen-Zhu · 2017
Cited alongside, same era.
Bridging the gap between constant step size stochastic gradient descent and markov chains
Original
A. Dieuleveut, A. Durmus, and F. Bach · 2017
Cited alongside, same era.
Stochastic modified equations and adaptive stochastic gradient algorithms
Q. Li, C. Tai, and E. Weinan · 2017
Cited alongside, same era.
Convergence rate of sign stochastic gradient descent for non-convex functions
J. Bernstein, K. Azizzadenesheli, Y.-X. Wang, and A. Anandkumar · 2018
Cited alongside, same era.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Cited alongside, same era.
The loss landscape of overparameterized neural networks
Original
Y. Cooper · 2018
Cited alongside, same era.
On the convergence of single-call stochastic extra-gradient methods
Original
Y.-G. Hsieh, F. Iutzeler, J. Malick, and P. Mertikopoulos · 2019
Cited alongside, same era.
Loss landscape sightseeing with multi-point optimization
Original
I. Skorokhodov and M. Burtsev · 2019
Cited alongside, same era.
Unified optimal analysis of the (stochastic) gradient method
Original
S. U. Stich · 2019
Cited alongside, same era.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
S. Vaswani, F. Bach, and M. Schmidt · 2019
Cited alongside, same era.