Fetching the paper…
Reading the bibliography…
Modern statistical inference tasks often require iterative optimization methods to compute the solution.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Stochastic estimation of the maximum of a regression function
J. Kiefer and J. Wolfowitz · 1952
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o (1/k2)
Y. Nesterov · 1983
Earlier work this paper cites.
Generalized linear models
P. McCullagh · 1984
Earlier work this paper cites.
Entropy and the central limit theorem
A. R. Barron · 1986
Earlier work this paper cites.
Efficient estimations from a slowly convergent Robbins–Monro process
D. Ruppert · 1988
Earlier work this paper cites.
New stochastic approximation type procedures
B. T. Polyak · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Natural gradient works efficiently in learning
S.-I. Amari · 1998
Earlier work this paper cites.
A strong approximation theorem for stochastic recursive algorithms
V. Borkar and S. Mitter · 1999
Earlier work this paper cites.
Numerical optimization
S. Wright and J. Nocedal · 1999
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
N. N. Schraudolph · 2002
Earlier work this paper cites.
Metastability in reversible diffusion processes i: Sharp asymptotics for capacities and exit times
A. Bovier, M. Eckhoff, V. Gayrard, and M. Klein · 2004
Earlier work this paper cites.
A stochastic quasi-newton method for online convex optimization
N. N. Schraudolph, J. Yu, and S. Günter · 2007
Earlier work this paper cites.
Self-normalized processes: Limit theory and Statistical Applications
V. H. Peña, T. L. Lai, and Q.-M. Shao · 2008
Earlier work this paper cites.
Sgd-qn: Careful quasi-newton stochastic gradient descent
A. Bordes, L. Bottou, and P. Gallinari · 2009
Earlier work this paper cites.
ggplot2: Elegant Graphics for Data Analysis
H. Wickham · 2009
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
R. Vershynin · 2010
Earlier work this paper cites.
Better mini-batch algorithms via accelerated gradient methods
A. Cotter, O. Shamir, N. Srebro, and K. Sridharan · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Bayesian learning via stochastic gradient langevin dynamics
M. Welling and Y. W. Teh · 2011
Cited alongside, same era.
Differential-geometrical methods in statistics , volume 28
S.-i. Amari · 2012
Cited alongside, same era.
A quasi-newton proximal splitting method
S. Becker and J. Fadili · 2012
Cited alongside, same era.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
S. Ghadimi and G. Lan · 2012
Cited alongside, same era.
Exact and inexact subsampled newton methods for optimization
R. Bollapragada, R. Byrd, and J. Nocedal · 2016
Later among the works it cites.
A stochastic quasi-newton method for large-scale optimization
R. H. Byrd, S. L. Hansen, J. Nocedal, and Y. Singer · 2016
Later among the works it cites.
Statistical inference for model parameters in stochastic gradient descent
X. Chen, J. D. Lee, X. T. Tong, and Y. Zhang · 2016
Later among the works it cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
S. Ghadimi and G. Lan · 2016
Later among the works it cites.
The landscape of empirical risk for non-convex losses
S. Mei, Y. Bai, and A. Montanari · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A stochastic gradient method with an exponential convergence _rate for finite training sets
N. L. Roux, M. Schmidt, and F. R. Bach · 2012
Cited alongside, same era.
Rate of convergence and edgeworth-type expansion in the entropic central limit theorem
S. G. Bobkov, G. P. Chistyakov, and F. Götze · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course , volume 87
Y. Nesterov · 2013
Cited alongside, same era.
Berry–esseen bounds in the entropic central limit theorem
S. G. Bobkov, G. P. Chistyakov, and F. Götze · 2014
Cited alongside, same era.
New insights and perspectives on the natural gradient method
J. Martens · 2014
Cited alongside, same era.
A linearly-convergent stochastic l-bfgs algorithm
P. Moritz, R. Nishihara, and M. Jordan · 2016
Later among the works it cites.
Sub-sampled newton methods with non-uniform sampling
P. Xu, J. Yang, F. Roosta-Khorasani, C. Ré, and M. W. Mahoney · 2016
Later among the works it cites.
Second-order stochastic optimization for machine learning in linear time
N. Agarwal, B. Bullins, and E. Hazan · 2017
Closest in time.
An investigation of newton-sketch and subsampled newton methods
A. S. Berahas, R. Bollapragada, and J. Nocedal · 2017
Closest in time.
Sampling from a log-concave distribution with compact support with proximal langevin monte carlo
N. Brosse, A. Durmus, É. Moulines, and M. Pereyra · 2017
Closest in time.
On variance reduction for stochastic smooth convex optimization with multiplicative noise
A. Jofré and P. Thompson · 2017
Closest in time.
Statistical inference using SGD
T. Li, L. Liu, A. Kyrillidis, and C. Caramanis · 2017
Closest in time.
Stochastic gradient descent as approximate bayesian inference
S. Mandt, M. D. Hoffman, and D. M. Blei · 2017
Closest in time.
Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis
M. Raginsky, A. Rakhlin, and M. Telgarsky · 2017
Closest in time.
Asymptotic and finite-sample properties of estimators based on stochastic gradients
P. Toulis, E. M. Airoldi, et al · 2017
Closest in time.
Stochastic quasi-newton methods for nonconvex stochastic optimization
X. Wang, S. Ma, D. Goldfarb, and W. Liu · 2017
Closest in time.
Efficient bayesian computation by proximal markov chain monte carlo: when langevin meets moreau
A. Durmus, E. Moulines, and M. Pereyra · 2018
Closest in time.
Local optimality and generalization guarantees for the langevin algorithm via empirical metastability
B. Tzen, T. Liang, and M. Raginsky · 2018
Closest in time.