Fetching the paper…
Reading the bibliography…
Many iterative procedures in stochastic optimization exhibit a transient phase followed by a stationary phase.
A stochastic approximation method
Robbins, H. and S. Monro (1951) · 1951
Earlier work this paper cites.
Accelerated stochastic approximation
Kesten, H. (1958) · 1958
Earlier work this paper cites.
Generalized linear models
McCullagh, P. (1984) · 1984
Earlier work this paper cites.
Numerical techniques for stochastic optimization
Ermoliev, Y. M. and R.-B. Wets (1988) · 1988
Earlier work this paper cites.
Stopping times for stochastic approximation
Yin, G. (1989) · 1989
Earlier work this paper cites.
Adaptive algorithms and stochastic approximations
Benveniste, A., P. Priouret, and M. Métivier (1990) · 1990
Earlier work this paper cites.
Non-asymptotic confidence bounds for stochastic approximation algorithms with constant step size
Pflug, G. C. (1990) · 1990
Earlier work this paper cites.
Gradient estimates for the performance of markov chains and discrete event processes
Pflug, G. C. (1992) · 1992
Earlier work this paper cites.
Accelerated stochastic approximation
Delyon, B. and A. Juditsky (1993) · 1993
Earlier work this paper cites.
A statistical study of on-line learning
Murata, N. (1998) · 1998
Earlier work this paper cites.
Beating the hold-out: Bounds for k-fold and progressive cross-validation
Blum, A., A. Kalai, and J. Langford (1999) · 1999
Earlier work this paper cites.
Stochastic approximation
Borkar, V. S. (2008) · 2008
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., A. Juditsky, G. Lan, and A. Shapiro (2009) · 2009
Cited alongside, same era.
Large-scale machine learning with stochastic gradient descent
Bottou, L. (2010) · 2010
Cited alongside, same era.
Implicit online learning
Kulis, B. and P. L. Bartlett (2010) · 2010
Cited alongside, same era.
Incremental proximal methods for large scale convex optimization
Bertsekas, D. P. (2011) · 2011
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, E. and F. R. Bach (2011) · 2011
Cited alongside, same era.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
Needell, D., R. Ward, and N. Srebro (2014) · 2014
Later among the works it cites.
Convergence of stochastic proximal gradient algorithm
Rosasco, L., S. Villa, and B. C. Vũ (2014) · 2014
Later among the works it cites.
Statistical analysis of stochastic gradient methods for generalized linear models
Toulis, P., J. Rennie, and E. Airoldi (2014) · 2014
Later among the works it cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
Zhang, T. (2004) · 2014
Later among the works it cites.
Scalable estimation strategies based on stochastic approximations: classical results and new insights
Toulis, P. and E. M. Airoldi (2015) · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xu, W. (2011) · 2011
Cited alongside, same era.
Stochastic Gradient Descent Tricks
Bottou, L. (2012) · 2012
Cited alongside, same era.
A stochastic gradient method with an exponential convergence _rate for finite training sets
Roux, N. L., M. Schmidt, and F. Bach (2012) · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and T. Zhang (2013) · 2013
Cited alongside, same era.
Proximal algorithms
Parikh, N. and S. Boyd (2013) · 2013
Cited alongside, same era.
Bottou, L., F. E. Curtis, and J. Nocedal (2016) · 2016
Later among the works it cites.
Towards stability and optimality in stochastic gradient descent
Toulis, P., D. Tran, and E. Airoldi (2016) · 2016
Later among the works it cites.
Asymptotic and finite-sample properties of estimators based on stochastic gradients
Toulis, P., E. M. Airoldi, et al. (2017) · 2017
Closest in time.
Gradient diversity: a key ingredient for scalable distributed learning
Yin, D., A. Pananjady, M. Lam, D. Papailiopoulos, K. Ramchandran, and P. Bartlett (2018) · 2018
Closest in time.
A proximal stochastic gradient method with progressive variance reduction
Xiao, L. and T. Zhang (2014) · 2075
Closest in time.