Fetching the paper…
Reading the bibliography…
Mini-batch stochastic gradient descent (SGD) and variants thereof approximate the objective function's gradient with a small number of training examples, aka the batch size.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Gradient methods for minimizing functionals
B. T. Polyak · 1963
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. S. Nemirovsky and D. B. Yudin · 1983
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Removing noise in on-line search using adaptive batch sizes
G. B. Orr · 1997
Earlier work this paper cites.
Online learning and stochastic approximations
L. Bottou · 1998
Earlier work this paper cites.
A statistical study of on-line learning
N. Murata · 1998
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Stochastic gradient descent tricks
L. Bottou · 2012
Earlier work this paper cites.
Sample size selection in optimization methods for machine learning
R. H. Byrd, G. M. Chin, J. Nocedal, and Y. Wu · 2012
Earlier work this paper cites.
Hybrid deterministic-stochastic methods for data fitting
M. P. Friedlander and M. Schmidt · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Y. Nesterov · 2013
Earlier work this paper cites.
No more pesky learning rates
T. Schaul, S. Zhang, and Y. LeCun · 2013
Earlier work this paper cites.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Convex optimization: Algorithms and complexity
S. Bubeck et al · 2015
Cited alongside, same era.
No more pesky learning rate guessing games
L. N. Smith · 2015
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina · 2016
Cited alongside, same era.
Big Batch SGD: Automated inference using adaptive batch sizes
S. De, A. Yadav, D. Jacobs, and T. Goldstein · 2016
Cited alongside, same era.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Later among the works it cites.
Deep learning tutorial at the Simons Institute, Berkeley, 2017, 2017
R. Salakhutdinov · 2017
Later among the works it cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi · 2017
Later among the works it cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
H. Xiao, K. Rasul, and R. Vollgraf · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Linear convergence of gradient and proximal-gradient methods under the Polyak-łojasiewicz condition
H. Karimi, J. Nutini, and M. Schmidt · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Cited alongside, same era.
Paleo: A performance model for deep neural networks
H. Qi, E. R. Sparks, and A. Talwalkar · 2016
Cited alongside, same era.
Imagenet pre-trained models with batch normalization
M. Simon, E. Rodner, and J. Denzler · 2016
Cited alongside, same era.
S. Zagoruyko and N. Komodakis · 2016
Cited alongside, same era.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic · 2017
Cited alongside, same era.
Coupling adaptive batch sizes with learning rates
L. Balles, J. Romero, and P. Hennig · 2017
Cited alongside, same era.
M. Belkin, D. J. Hsu, and P. Mitra · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Later among the works it cites.
Stability and generalization of learning algorithms that converge to global optima
Z. Charles and D. Papailiopoulos · 2018
Later among the works it cites.
A bayesian perspective on generalization and stochastic gradient descent
S. L. Smith and Q. V. Le · 2018
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
S. L. Smith, P.-J. Kindermans, and Q. V. Le · 2018
Later among the works it cites.
Gradient diversity: a key ingredient for scalable distributed learning
D. Yin, A. Pananjady, M. Lam, D. Papailiopoulos, K. Ramchandran, and P. Bartlett · 2018
Later among the works it cites.
New insight into hybrid stochastic gradient descent: Beyond with-replacement sampling and convexity
P. Zhou, X. Yuan, and J. Feng · 2018
Later among the works it cites.
The youtube-8m kaggle competition: challenges and methods
H. Zou, K. Xu, J. Li, and J. Zhu · 2018
Later among the works it cites.
History-gradient aided batch size adaptation for variance reduced algorithms
K. Ji, Z. Wang, B. Weng, Y. Zhou, W. Zhang, and Y. Liang · 2019
Closest in time.
Tight dimension independent lower bound on the expected convergence rate for diminishing step sizes in sgd
P. H. Nguyen, L. Nguyen, and M. van Dijk · 2019
Closest in time.
Painless stochastic gradient: Interpolation, line-search, and convergence rates
S. Vaswani, A. Mishkin, I. Laradji, M. Schmidt, G. Gidel, and S. Lacoste-Julien · 2019
Closest in time.
AdaGrad stepsizes: Sharp convergence over nonconvex landscapes
R. Ward, X. Wu, and L. Bottou · 2019
Closest in time.