Fetching the paper…
Reading the bibliography…
In many applications involving large dataset or online updating, stochastic gradient descent (SGD) provides a scalable way to compute parameter estimates and has gained increasing popularity due to its numerical convenience and memory efficiency.
Robbins, H. & Monro, S. (1951), ‘A stochastic approximation method’, The annals of mathematical statistics pp. 400–407
1951
Earlier work this paper cites.
Nelder, J. A. & Baker, R. J. (1972), ‘Generalized linear models’, Encyclopedia of statistical sciences
1972
Earlier work this paper cites.
Rubin, D. B. et al. (1981), ‘The bayesian bootstrap’, The annals of statistics
1981
Earlier work this paper cites.
Ruppert, D. (1988), Efficient estimations from a slowly convergent robbins-monro process, Technical report, Cornell University Operations Research and Industrial Engineering
1988
Earlier work this paper cites.
Polyak, B. T. & Juditsky, A. B. (1992), ‘Acceleration of stochastic approximation by averaging’, SIAM Journal on Control and Optimization
1992
Cited alongside, same era.
Rao, C. R. & Zhao, L. (1992), ‘Approximation to the distribution of m-estimates in linear models by randomly weighted bootstrap’, Sankhyā: The Indian Journal of Statistics, Series A pp. 323–331
1992
Cited alongside, same era.
Hastie, T., Tibshirani, R. & Friedman, J. (2009), ‘The elements of statistical learning 2nd edition’
2009
Cited alongside, same era.
Shao, J. & Tu, D. (2012), The jackknife and bootstrap, Springer Science & Business Media
2012
Cited alongside, same era.
2014
Later among the works it cites.
2015
Later among the works it cites.
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…