Fetching the paper…
Reading the bibliography…
Stopping criteria for Stochastic Gradient Descent (SGD) methods play important roles from enabling adaptive step size schemes to providing rigor for downstream analyses such as asymptotic inference.
The annals of mathematical statistics pp. 400–407 (1951)
Robbins, H., Monro, S.: A stochastic approximation method · 1951
Earlier work this paper cites.
The Annals of Mathematical Statistics 23
Kiefer, J., Wolfowitz, J., et al.: Stochastic estimation of the maximum of a regression function · 1952
Earlier work this paper cites.
The Annals of Mathematical Statistics 25
Chung, K.L., et al.: On a stochastic approximation method · 1954
Earlier work this paper cites.
The Annals of Mathematical Statistics pp. 237–247 (1962)
Farrell, R.: Bounded length confidence intervals for the zero of a regression function · 1962
Earlier work this paper cites.
The Annals of Mathematical Statistics pp. 191–200 (1967)
Fabian, V.: Stochastic approximation of minima with improved asymptotic speed · 1967
Earlier work this paper cites.
Integer and nonlinear programming pp. 37–86 (1970)
Zoutendijk, G.: Nonlinear programming, computational methods · 1970
Earlier work this paper cites.
Zeitschrift für Wahrscheinlichkeitstheorie und verwandte Gebiete 26
Sielken, R.L.: Stopping times for stochastic approximation procedures · 1973
Earlier work this paper cites.
Zeitschrift für Wahrscheinlichkeitstheorie und verwandte Gebiete 60
Stroup, D.F., Braun, H.I.: On a new stopping rule for stochastic approximation · 1982
Earlier work this paper cites.
Stochastics: An International Journal of Probability and Stochastic Processes 9
Ermoliev, Y.: Stochastic quasigradient methods and their application to system optimization · 1983
Earlier work this paper cites.
Zhurnal Vychislitel’noi Matematiki i Matematicheskoi Fiziki 23
Mirozahmedov, F., Uryasev, S.: Adaptive stepsize regulation for stochastic optimization algorithm · 1983
Earlier work this paper cites.
numerical techniques for stochastic optimization pp. 353–372 (1988)
Pflug, G.C.: Stepsize rules, stopping times and their implementation in stochastic quasi-gradient algorithms · 1988
Earlier work this paper cites.
Journal of Optimization Theory and Applications 67
Yin, G.: A stopping rule for the robbins-monro method · 1990
Earlier work this paper cites.
In: Probabilistic methods for algorithmic discrete mathematics, pp. 195–248. Springer (1998)
McDiarmid, C.: Concentration · 1998
Earlier work this paper cites.
In: Neural Networks: Tricks of the trade, pp. 55–69. Springer (1998)
Prechelt, L.: Early stopping-but when? · 1998
Earlier work this paper cites.
Cambridge university press (2000)
Van der Vaart, A.W.: Asymptotic statistics, vol. 3 · 2000
Earlier work this paper cites.
SIAM Journal on optimization 19
Nemirovski, A., Juditsky, A., Lan, G., Shapiro, A.: Robust stochastic approximation approach to stochastic programming · 2009
Earlier work this paper cites.
Chapman and Hall/CRC (2009)
Wu, L.: Mixed effects models for complex data · 2009
Earlier work this paper cites.
Cambridge university press (2010)
Durrett, R.: Probability: theory and examples, 4th edn · 2010
Earlier work this paper cites.
In: 49th IEEE Conference on Decision and Control (CDC), pp. 4171–4176. IEEE (2010)
Wada, T., Itani, T., Fujisaki, Y.: A stopping rule for linear stochastic approximation · 2010
Earlier work this paper cites.
Optimization for Machine Learning 2010
Bertsekas, D.P.: Incremental gradient, subgradient, and proximal methods for convex optimization: A survey · 2011
Cited alongside, same era.
Springer Science & Business Media (2013)
Devroye, L., Györfi, L., Lugosi, G.: A probabilistic theory of pattern recognition, vol. 31 · 2013
Cited alongside, same era.
Mathematical Programming 155
Ghadimi, S., Lan, G., Zhang, H.: Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization · 2016
Cited alongside, same era.
In: Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pp. 795–811. Springer (2016)
Karimi, H., Nutini, J., Schmidt, M.: Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition · 2016
Cited alongside, same era.
SIAM Journal on Optimization 26
Patel, V.: Kalman-based stochastic gradient method with stop condition and insensitivity to conditioning · 2016
Cited alongside, same era.
In: International conference on machine learning, pp. 314–323 (2016)
In: Advances in Neural Information Processing Systems, pp. 5564–5574 (2018)
Li, Z., Li, J.: A simple proximal stochastic gradient method for nonsmooth nonconvex optimization · 2018
Later among the works it cites.
arXiv preprint arXiv:1806.01811 (2018)
Ward, R., Wu, X., Bottou, L.: Adagrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization · 2018
Later among the works it cites.
Ph.D. thesis, The Ohio State University (2018)
Zhou, Y.: Nonconvex optimization in machine learning: Convergence, landscape, and generalization · 2018
Later among the works it cites.
In: Pacific Rim International Conference on Artificial Intelligence, pp. 337–349. Springer (2019)
Bi, J., Gunn, S.R.: A stochastic gradient method with biased estimation for faster nonconvex optimization · 2019
Later among the works it cites.
arXiv preprint arXiv:1902.00247 (2019)
Fang, C., Lin, Z., Zhang, T.: Sharp analysis for nonconvex sgd escaping from saddle points · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reddi, S.J., Hefny, A., Sra, S., Poczos, B., Smola, A.: Stochastic variance reduction for nonconvex optimization · 2016
Cited alongside, same era.
In: Advances in Neural Information Processing Systems, pp. 1145–1153 (2016)
Reddi, S.J., Sra, S., Poczos, B., Smola, A.J.: Proximal stochastic methods for nonsmooth nonconvex finite-sum optimization · 2016
Cited alongside, same era.
arXiv preprint arXiv:1705.07562 (2017)
Hu, W., Li, C.J., Li, L., Liu, J.G.: On the diffusion approximation of nonconvex stochastic gradient descent · 2017
Cited alongside, same era.
arXiv preprint arXiv:1704.07953 (2017)
Huang, F., Chen, S.: Linear convergence of accelerated stochastic gradient descent for nonconvex nonsmooth optimization · 2017
Cited alongside, same era.
Ma, Y., Klabjan, D.: Convergence analysis of batch normalization for deep neural nets · 2017
Cited alongside, same era.
arXiv preprint arXiv:1709.04718 9
Patel, V.: The impact of local geometry and batch size on the convergence and divergence of stochastic gradient descent · 2017
Cited alongside, same era.
Xiao, H., Rasul, K., Vollgraf, R.: Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms (2017)
2017
Cited alongside, same era.
Later among the works it cites.
arXiv preprint arXiv:1904.01517 (2019)
Fehrman, B., Gess, B., Jentzen, A.: Convergence rates for the stochastic gradient descent method for non-convex objective functions · 2019
Later among the works it cites.
arXiv preprint arXiv:1902.04811 (2019)
Jin, C., Netrapalli, P., Ge, R., Kakade, S.M., Jordan, M.I.: On nonconvex optimization for machine learning: Gradients, stochasticity, and saddle points · 2019
Later among the works it cites.
IEEE Transactions on Neural Networks and Learning Systems (2019)
Lei, Y., Hu, T., Li, G., Tang, K.: Stochastic gradient descent for nonconvex learning without bounded gradient assumptions · 2019
Later among the works it cites.
arXiv preprint arXiv:1906.11417 (2019)
Park, S., Jung, S.H., Pardalos, P.M.: Combining stochastic adaptive cubic regularization with negative curvature for nonconvex optimization · 2019
Later among the works it cites.
Optimization Methods and Software 34
Wang, X., Wang, X., Yuan, Y.x.: Stochastic proximal quasi-newton methods for non-convex composite optimization · 2019
Later among the works it cites.
arXiv preprint arXiv:1905.04346 (2019)
Yu, H., Jin, R.: On the computation and communication complexity of parallel sgd with dynamic batch sizes for stochastic non-convex optimization · 2019
Later among the works it cites.
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 11127–11135 (2019)
Zou, F., Shen, L., Jie, Z., Zhang, W., Liu, W.: A sufficient condition for convergences of adam and rmsprop · 2019
Later among the works it cites.
arXiv preprint arXiv:2001.06699 (2020)
Curtis, F.E., Scheinberg, K.: Adaptive stochastic optimization · 2020
Closest in time.
arXiv preprint arXiv:2006.10311 (2020)
Gower, R.M., Sebbouh, O., Loizou, N.: Sgd for structured nonconvex functions: Learning rates, minibatching and interpolation · 2020
Closest in time.
arXiv preprint arXiv:2002.03329 (2020)
Khaled, A., Richtárik, P.: Better theory for sgd in the nonconvex world · 2020
Closest in time.
arXiv preprint arXiv:2006.11144 (2020)
Mertikopoulos, P., Hallak, N., Kavis, A., Cevher, V.: On the almost sure convergence of stochastic gradient descent in non-convex problems · 2020
Closest in time.
Annual Review of Statistics and Its Application 7
Roy, V.: Convergence diagnostics for markov chain monte carlo · 2020
Closest in time.
arXiv preprint arXiv:2002.10597 (2020)
Zhang, P., Lang, H., Liu, Q., Xiao, L.: Statistical adaptive stochastic gradient methods · 2020
Closest in time.