Fetching the paper…
Reading the bibliography…
The stochastic gradient descent (SGD) algorithm is widely used for parameter estimation, especially for huge data sets and online learning.
A generalization of regularized dual averaging and its dynamics
Chao, S.-K. and Cheng, G. (2019) · 1909
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S. (1951) · 1951
Earlier work this paper cites.
Stochastic estimation of the maximum of a regression function
Kiefer, J. and Wolfowitz, J. (1952) · 1952
Earlier work this paper cites.
Approximation methods which converge with probability one
Blum, J. R. (1954) · 1954
Earlier work this paper cites.
On stochastic approximation
Dvoretzky, A. (1956) · 1956
Earlier work this paper cites.
Asymptotic distribution of stochastic approximation procedures
Sacks, J. (1958) · 1958
Earlier work this paper cites.
On asymptotic normality in stochastic approximation
Fabian, V. (1968) · 1968
Earlier work this paper cites.
A convergence theorem for non negative almost supermartingales and some applications
Robbins, H. and Siegmund, D. (1971) · 1971
Earlier work this paper cites.
Analysis of recursive stochastic algorithms
Ljung, L. (1977) · 1977
Earlier work this paper cites.
Sharp inequalities for martingales and stochastic integrals
Burkholder, D. L. (1988) · 1988
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
Ruppert, D. (1988) · 1988
Earlier work this paper cites.
Estimating the asymptotic variance with batch means
Glynn, P. W. and Whitt, W. (1991) · 1991
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B. (1992) · 1992
Earlier work this paper cites.
Online learning and stochastic approximations
Bottou, L. (1998) · 1998
Earlier work this paper cites.
Theoretical comparisons of block bootstrap methods
Lahiri, S. N. (1999) · 1999
Earlier work this paper cites.
Subsampling
Politis, D. N., Romano, J. P., and Wolf, M. (1999) · 1999
Cited alongside, same era.
Resampling methods for dependent data
Lahiri, S. N. (2003) · 2003
Cited alongside, same era.
Stochastic approximation
Lai, T. L. (2003) · 2003
Cited alongside, same era.
Fixed-width output analysis for markov chain monte carlo
Jones, G. L., Haran, M., Caffo, B. S., and Neath, R. (2006) · 2006
Cited alongside, same era.
Computer-intensive rate estimation, diverging statistics and scanning
McElroy, T., Politis, D. N., et al. (2007) · 2007
Cited alongside, same era.
Predicting clicks: estimating the click-through rate for new ads
Richardson, M., Dominowska, E., and Ragno, R. (2007) · 2007
Cited alongside, same era.
A progressive block empirical likelihood method for time series
Kim, Y. M., Lahiri, S. N., and Nordman, D. J. (2013) · 2013
Later among the works it cites.
A nonstandard empirical likelihood for time series
Nordman, D. J., Bunzel, H., and Lahiri, S. N. (2013) · 2013
Later among the works it cites.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Shamir, O. and Zhang, T. (2013) · 2013
Later among the works it cites.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
Hazan, E. and Kale, S. (2014) · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Later among the works it cites.
Bid-aware gradient descent for unbiased learning with censored data in display advertising
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rio, E. (2009) · 2009
Cited alongside, same era.
Recursive estimation of time-average variance constants
Wu, W. B. (2009) · 2009
Cited alongside, same era.
Batch means and spectral variance estimators in Markov chain Monte Carlo
Flegal, J. M. and Jones, G. L. (2010) · 2010
Cited alongside, same era.
Online learning for latent dirichlet allocation
Hoffman, M., Bach, F. R., and Blei, D. M. (2010) · 2010
Cited alongside, same era.
Online learning for matrix factorization and sparse coding
Mairal, J., Bach, F. R., Ponce, J., and Sapiro, G. (2010) · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Cited alongside, same era.
Zhang, W., Zhou, T., Wang, J., and Xu, J. (2016) · 2016
Later among the works it cites.
Asymptotic and finite-sample properties of estimators based on stochastic gradients
Toulis, P. and Airoldi, E. M. (2017) · 2017
Later among the works it cites.
Online bootstrap confidence intervals for the stochastic gradient descent estimator
Fang, Y., Xu, J., and Yang, L. (2018) · 2018
Later among the works it cites.
Su, W. and Zhu, Y. (2018) · 2018
Later among the works it cites.
Scalable statistical inference for averaged implicit stochastic gradient descent
Fang, Y. (2019) · 2019
Later among the works it cites.
Statistical inference for the population landscape via moment-adjusted stochastic gradients
Liang, T. and Su, W. (2019) · 2019
Later among the works it cites.
Multivariate output analysis for markov chain monte carlo
Vats, D., Flegal, J. M., and Jones, G. L. (2019) · 2019
Later among the works it cites.
Statistical inference for model parameters in stochastic gradient descent
Chen, X., Lee, J. D., Tong, X. T., and Zhang, Y. (2020) · 2020
Closest in time.
On linear stochastic approximation: Fine-grained polyak-ruppert and non-asymptotic concentration
Mou, W., Li, C. J., Wainwright, M. J., Bartlett, P. L., and Jordan, M. I. (2020) · 2020
Closest in time.
Empirical likelihood methods with weakly dependent processes
Kitamura, Y. et al. (1997) · 2084
Closest in time.