Fetching the paper…
Reading the bibliography…
Increasing the batch size is a popular way to speed up neural network training, but beyond some critical batch size, larger batch sizes yield diminishing returns.
Fast convergence of natural gradient descent for overparameterized neural networks
Guodong Zhang, James Martens, and Roger Grosse · 1905
Earlier work this paper cites.
Fundamental Methods of Mathematical Economics
A.C. Chiang · 1974
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Optimal stochastic search and adaptive momentum
Todd K. Leen and Genevieve B. Orr · 1994
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N Schraudolph · 2002
Earlier work this paper cites.
The tradeoffs of large scale learning
Léon Bottou and Olivier Bousquet · 2008
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate o (1/n)
Francis Bach and Eric Moulines · 2013
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
New insights and perspectives on the natural gradient method
James Martens · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Three mechanisms of weight decay regularization
Guodong Zhang, Chaoqi Wang, Bowen Xu, and Roger Grosse · 2015
Earlier work this paper cites.
Nonparametric stochastic approximation with large step-sizes
Aymeric Dieuleveut and Francis Bach · 2016
Earlier work this paper cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Earlier work this paper cites.
A kronecker-factored approximate fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
On the influence of momentum acceleration on online learning
Kun Yuan, Bicheng Ying, and Ali H. Sayed · 2016
Cited alongside, same era.
Distributed second-order optimization using Kronecker-factored approximations
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Parallelizing stochastic gradient descent for least squares regression: mini-batching, averaging, and model misspecification
Prateek Jain, Sham M Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2018
Later among the works it cites.
Three factors influencing minima in SGD
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2018
Later among the works it cites.
On the insufficiency of existing momentum schemes for stochastic optimization
Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham Kakade · 2018
Later among the works it cites.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jimmy Ba, Roger Grosse, and James Martens · 2017
Cited alongside, same era.
Coupling adaptive batch sizes with learning rates
Lukas Balles, Javier Romero, and Philipp Hennig · 2017
Cited alongside, same era.
Critical hyper-parameters: No random, no cry
Olivier Bousquet, Sylvain Gelly, Karol Kurach, Olivier Teytaud, and Damien Vincent · 2017
Cited alongside, same era.
Why momentum really works
Gabriel Goh · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: Training Imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Train longer, generalize better: Closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Fast estimation of tr(f(a)) via stochastic lanczos quadrature
Shashanka Ubaru, Jie Chen, and Yousef Saad · 2017
Cited alongside, same era.
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Later among the works it cites.
Second-order optimization method for large mini-batch: Training resnet-50 on imagenet in 35 epochs
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2018
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
Christopher J Shallue, Jaehoon Lee, Joe Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl · 2018
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Samuel L. Smith, Pieter-Jan Kindermans, and Quoc V. Le · 2018
Later among the works it cites.
Understanding short-horizon bias in stochastic meta-optimization
Yuhuai Wu, Mengye Ren, Renjie Liao, and Roger Grosse · 2018
Later among the works it cites.
The physical systems behind optimization algorithms
Lin Yang, Raman Arora, Tuo Zhao, et al · 2018
Later among the works it cites.
Gradient diversity: a key ingredient for scalable distributed learning
Dong Yin, Ashwin Pananjady, Max Lam, Dimitris Papailiopoulos, Kannan Ramchandran, and Peter Bartlett · 2018
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Closest in time.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Closest in time.
Momentum enables large batch training
Samuel L Smith, Erich Elsen, and Soham De · 2019
Closest in time.
Eigendamage: Structured pruning in the Kronecker-factored eigenbasis
Chaoqi Wang, Roger Grosse, Sanja Fidler, and Guodong Zhang · 2019
Closest in time.