Fetching the paper…
Reading the bibliography…
Recent hardware developments have dramatically increased the scale of data parallelism available for neural network training.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Stochastic estimation of the maximum of a regression function
Jack Kiefer, Jacob Wolfowitz, et al · 1952
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence O ( 1 / k 2 ) {O}(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1986
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Efficient backprop
Yann Le Cun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 1998
Earlier work this paper cites.
MNIST handwritten digit database, 1998
Yann LeCun, Corinna Cortes, and CJ Burges · 1998
Earlier work this paper cites.
The general inefficiency of batch training for gradient descent learning
D Randall Wilson and Tony R Martinez · 2003
Earlier work this paper cites.
The tradeoffs of large scale learning
Léon Bottou and Olivier Bousquet · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Neural networks for machine learning, lecture 6a: overview of mini-batch gradient descent, 2012
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
An optimal method for stochastic composite optimization
Guanghui Lan · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George E. Dahl, and Geoffrey E. Hinton · 2013
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson · 2014
Earlier work this paper cites.
Efficient mini-batch training for stochastic optimization
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J Smola · 2014
Earlier work this paper cites.
Understanding machine learning: From foundations to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
ImageNet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Cited alongside, same era.
TensorFlow: a system for large-scale machine learning
Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Revisiting distributed synchronous SGD
Jianmin Chen, Xinghao Pan, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Cited alongside, same era.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit
Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, et al · 2017
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Later among the works it cites.
OpenImages: A public dataset for large-scale multi-label and multi-class image classification., 2017
Ivan Krasin, Tom Duerig, Neil Alldrin, Vittorio Ferrari, Sami Abu-El-Haija, Alina Kuznetsova, Hassan Rom, Jasper Uijlings, Stefan Popov, Shahab Kamali, Matteo Malloci, Jordi Pont-Tuset, Andreas Veit, Serge Belongie, Victor Gomes, Abhinav Gupta, Chen Sun, Gal Chechik, David Cai, Zheyun Feng, Dhyanesh Narayanan, and Kevin Murphy · 2017
Later among the works it cites.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A Kronecker-factored approximate Fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Cited alongside, same era.
Introduction to online convex optimization
Elad Hazan · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu · 2016
Cited alongside, same era.
Without-replacement sampling for stochastic gradient methods
Ohad Shamir · 2016
Cited alongside, same era.
Rethinking the Inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Extremely large minibatch SGD: Training ResNet-50 on ImageNet in 15 minutes
Takuya Akiba, Shuji Suzuki, and Keisuke Fukuda · 2017
Cited alongside, same era.
Later among the works it cites.
Learning from noisy large-scale datasets with minimal supervision
Andreas Veit, Neil Alldrin, Gal Chechik, Ivan Krasin, Abhinav Gupta, and Serge J Belongie · 2017
Later among the works it cites.
Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Later among the works it cites.
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer · 2017
Later among the works it cites.
Large scale distributed neural network training through online distillation
Rohan Anil, Gabriel Pereyra, Alexandre Passos, Robert Ormandi, George E. Dahl, and Geoffrey E. Hinton · 2018
Closest in time.
The effect of network width on the performance of large-batch training
Lingjiao Chen, Hongyi Wang, Jinman Zhao, Dimitris Papailiopoulos, and Paraschos Koutris · 2018
Closest in time.
On the computational inefficiency of large batch sizes for stochastic gradient descent
Noah Golmant, Nikita Vemuri, Zhewei Yao, Vladimir Feinberg, Amir Gholami, Kai Rothauge, Michael W Mahoney, and Joseph Gonzalez · 2018
Closest in time.
Parallelizing stochastic gradient descent for least squares regression: Mini-batching, averaging, and model misspecification
Prateek Jain, Sham M. Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2018
Closest in time.
Universal statistics of Fisher information in deep neural networks: Mean field approach
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 2018
Closest in time.
On the insufficiency of existing momentum schemes for stochastic optimization
Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham M. Kakade · 2018
Closest in time.
Don’t use large mini-batches, use local SGD
Tao Lin, Sebastian U Stich, and Martin Jaggi · 2018
Closest in time.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Closest in time.
Revisiting small batch training for deep neural networks
Dominic Masters and Carlo Luschi · 2018
Closest in time.
A Bayesian perspective on generalization and stochastic gradient descent
Samuel L. Smith and Quoc V. Le · 2018
Closest in time.
Understanding short-horizon bias in stochastic meta-optimization
Yuhuai Wu, Mengye Ren, Renjie Liao, and Roger Grosse · 2018
Closest in time.
Gradient diversity: a key ingredient for scalable distributed learning
Dong Yin, Ashwin Pananjady, Max Lam, Dimitris Papailiopoulos, Kannan Ramchandran, and Peter Bartlett · 2018
Closest in time.