Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) is a popular and efficient method with wide applications in training deep neural nets and other nonconvex models.
B. T. Polyak, “Gradient methods for minimizing functionals,”
1963
Earlier work this paper cites.
J. L. Doob,
1994
Earlier work this paper cites.
D. P. Bertsekas and J. N. Tsitsiklis, “Gradient convergence in gradient methods with errors,”
2000
Earlier work this paper cites.
T. Zhang, “Solving large scale linear prediction problems using stochastic gradient descent algorithms,” in
2004
Earlier work this paper cites.
Y. Ying and D.-X. Zhou, “Online regularized classification algorithms,”
2006
Earlier work this paper cites.
E. Hazan, A. Agarwal, and S. Kale, “Logarithmic regret algorithms for online convex optimization,”
2007
Earlier work this paper cites.
T. Hu and D.-X. Zhou, “Online learning with samples drawn from non-identical distributions,”
2009
Earlier work this paper cites.
S. Ghadimi and G. Lan, “Stochastic first-and zeroth-order methods for nonconvex stochastic programming,”
2013
Earlier work this paper cites.
H. Sun and Q. Wu, “Sparse representation in kernel machines,”
2015
Cited alongside, same era.
H. Karimi, J. Nutini, and M. Schmidt, “Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition,” in
2016
Cited alongside, same era.
S. Reddi, A. Hefny, S. Sra, B. Poczos, and A. Smola, “Stochastic variance reduction for nonconvex optimization,” in
2016
Cited alongside, same era.
S. Ghadimi, G. Lan, and H. Zhang, “Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization,”
2016
Cited alongside, same era.
Y. Ying and D.-X. Zhou, “Unregularized online learning algorithms with general loss functions,”
2017
Cited alongside, same era.
L. Bottou, F. E. Curtis, and J. Nocedal, “Optimization methods for large-scale machine learning,”
2018
Later among the works it cites.
J. Lin and D.-X. Zhou, “Online learning algorithms can converge comparably fast as batch learning,”
2018
Later among the works it cites.
Y. Lei and K. Tang, “Stochastic composite mirror descent: Optimal bounds with high probabilities,” in
2018
Later among the works it cites.
Y. Lei and D.-X. Zhou, “Convergence of online mirror descent,”
2018
Later among the works it cites.
Y. Lei, L. Shi, and Z.-C. Guo, “Convergence of unregularized online learning algorithms,”
2018
Later among the works it cites.
L. M. Nguyen, P. H. Nguyen, M. van Dijk, P. Richtárik, K. Scheinberg, and M. Takáč, “SGD and hogwild! convergence without the bounded gradients assumption,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Chang, M. Lin, and C. Zhang, “On the generalization ability of online gradient descent algorithm under the quadratic growth condition,”
2018
Cited alongside, same era.
D. J. Foster, A. Sekhari, and K. Sridharan, “Uniform convergence of gradients for non-convex learning and optimization,” in
2018
Cited alongside, same era.
2018
Later among the works it cites.
S.-B. Lin and D.-X. Zhou, “Distributed kernel-based gradient descent algorithms,”
2018
Later among the works it cites.