Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J. (2018) · 2018
Later among the works it cites.
Gradient descent learns linear dynamical systems
Hardt, M., Ma, T., and Recht, B. (2018) · 2018
Later among the works it cites.
An alternative view: When does SGD escape local minima?
Kleinberg, B., Li, Y., and Yuan, Y. (2018) · 2018
Later among the works it cites.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Ma, S., Bassily, R., and Belkin, M. (2018) · 2018
Later among the works it cites.
Linear convergence of first order methods for non-strongly convex optimization
Necoara, I., Nesterov, Y., and Glineur, F. (2018) · 2018
Later among the works it cites.
SGD and hogwild! Convergence without the bounded gradients assumption
Nguyen, L., Nguyen, P. H., van Dijk, M., Richtárik, P., Scheinberg, K., and Takáč, M. (2018) · 2018
Later among the works it cites.
L4: Practical loss-based stepsize adaptation for deep learning
Rolinek, M. and Martius, G. (2018) · 2018
Later among the works it cites.
SGD: General analysis and improved rates
Gower, R. M., Loizou, N., Qian, X., Sailanbayev, A., Shulgin, E., and Richtárik, P. (2019) · 2019
Later among the works it cites.
Stochastic gradient descent for nonconvex learning without bounded gradient assumptions
Lei, Y., Hu, T., Li, G., and Tang, K. (2019) · 2019
Later among the works it cites.
Randomized iterative methods for linear systems: momentum, inexactness and gossip
Loizou, N. (2019) · 2019
Later among the works it cites.
Towards closing the gap between the theory and practice of SVRG
Sebbouh, O., Gazagnadou, N., Jelassi, S., Bach, F., and Gower, R. (2019) · 2019
Later among the works it cites.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes
Ward, R., Wu, X., and Bottou, L. (2019) · 2019
Later among the works it cites.
SGD converges to global minimum in deep learning via star-convex path
Zhou, Y., Yang, J., Zhang, H., Liang, Y., and Tarokh, V. (2019) · 2019
Later among the works it cites.
Stochastic quasi-gradient methods: Variance reduction via jacobian sketching
Gower, R. M., Richtárik, P., and Bach, F. (2020) · 2020
Closest in time.
Near-optimal methods for minimizing star-convex functions and beyond
Hinder, O., Sidford, A., and Sohoni, N. (2020) · 2020
Closest in time.
A unified theory of decentralized SGD with changing topology and local updates
Koloskova, A., Loizou, N., Boreiri, S., Jaggi, M., and Stich, S. U. (2020) · 2020
Closest in time.
On the variance of the adaptive learning rate and beyond
Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J. (2020) · 2020
Closest in time.
Stochastic reformulations of linear systems: algorithms and convergence theory
Richtárik, P. and Takác, M. (2020) · 2020
Closest in time.