Fetching the paper…
Reading the bibliography…
In this note we give a simple proof for the convergence of stochastic gradient (SGD) methods on $\mu$-convex functions under a (milder than standard) $L$-smoothness assumption.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
On a stochastic approximation method
K. L. Chung · 1954
Earlier work this paper cites.
Efficient estimations from a slowly convergent Robbins-Monro process
D. Ruppert · 1988
Earlier work this paper cites.
New method of stochastic approximation type
B. T. Polyak · 1990
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
Introductory Lectures on Convex Optimization , volume 87 of Springer Science & Business Media
Yurii Nesterov · 2004
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Francis R. Bach and Eric Moulines · 2011
Earlier work this paper cites.
Simon Lacoste-Julien, Mark W. Schmidt, and Francis R. Bach · 2012
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Cited alongside, same era.
Fast convergence of stochastic gradient descent under a strong growth condition
Mark Schmidt and Nicolas Le Roux · 2013
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Ohad Shamir and Tong Zhang · 2013
Cited alongside, same era.
First-order methods of smooth convex optimization with inexact oracle
Olivier Devolder, François Glineur, and Yurii Nesterov · 2014
Cited alongside, same era.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Later among the works it cites.
SGD and Hogwild! Convergence without the bounded gradients assumption
Lam Nguyen, Phuong Ha Nguyen, Marten van Dijk, Peter Richtárik, Katya Scheinberg, and Martin Takáč · 2018
Later among the works it cites.
Sparsified SGD with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Later among the works it cites.
SGD: General analysis and improved rates
Robert M. Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Closest in time.
Convergence rates for deterministic and stochastic subgradient methods without lipschitz continuity
B. Grimmer · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
Deanna Needell, Nathan Srebro, and Rachel Ward · 2016
Cited alongside, same era.
Optimization methods for large-scale machine learning
L. Bottou, F. Curtis, and J. Nocedal · 2018
Cited alongside, same era.
Stochastic quasi-gradient methods: Variance reduction via jacobian sketching
Robert M. Gower, Peter Richtárik, and Francis Bach · 2018
Cited alongside, same era.
Sai P. Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi, Sebastian U. Stich, and Ananda T. Suresh · 2019
Closest in time.
Linear convergence of first order methods for non-strongly convex optimization
I. Necoara, Yu. Nesterov, and F. Glineur · 2019
Closest in time.