Fetching the paper…
Reading the bibliography…
Large-batch stochastic gradient descent (SGD) is widely used for training in distributed deep learning because of its training-time efficiency, however, extremely large-batch SGD leads to poor generalization and easily converges to sharp minima, which prevents naive large-scale data-parallel SGD (DP-SGD) from converging to good minima.
Yeming Wen, Kevin Luk, Maxime Gazeau, Guodong Zhang, Harris Chan, and Jimmy Ba · 1902
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Firecaffe: near-linear acceleration of deep neural network training on compute clusters
Forrest N. Iandola, Khalid Ashraf, Matthew W. Moskewicz, and Kurt Keutzer · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Chainer : a next-generation open source framework for deep learning
Seiya Tokui and Kenta Oono · 2015
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer T. Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
ChainerMN: Scalable Distributed Deep Learning Framework
Takuya Akiba, Keisuke Fukuda, and Shuji Suzuki · 2017
Earlier work this paper cites.
Extremely large minibatch SGD: training resnet-50 on imagenet in 15 minutes
Takuya Akiba, Shuji Suzuki, and Keisuke Fukuda · 2017
Cited alongside, same era.
Pratik Chaudhari and Stefano Soatto · 2017
Cited alongside, same era.
Valeriu Codreanu, Damian Podareanu, and Vikram Saletore · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Xianyan Jia, Shutao Song, Wei He, Yangzihao Wang, Haidong Rong, Feihu Zhou, Liqiang Xie, Zhenyu Guo, Yuanzhou Yang, Liwei Yu, Tiegang Chen, Guangxiao Hu, Shaohuai Shi, and Xiaowen Chu · 2018
Later among the works it cites.
An alternative view: When does SGD escape local minima?
Bobby Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Later among the works it cites.
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2018
Later among the works it cites.
How Does Batch Normalization Help Optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Later among the works it cites.
Smoothout: Smoothing out sharp minima for generalization in large-batch deep learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hao Li, Zheng Xu, Gavin Taylor, and Tom Goldstein · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V. Ugur Güney, Yann Dauphin, and Léon Bottou · 2017
Cited alongside, same era.
Large Batch Training of Convolutional Networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
Random erasing data augmentation
Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang · 2017
Cited alongside, same era.
Escaping saddles with stochastic gradients
Hadi Daneshmand, Jonas Kohler, Aurelien Lucchi, and Thomas Hofmann · 2018
Cited alongside, same era.
Wei Wen, Yandan Wang, Feng Yan, Cong Xu, Yiran Chen, and Hai Li · 2018
Later among the works it cites.
Image classification at supercomputer scale
Chris Ying, Sameer Kumar, Dehao Chen, Tao Wang, and Youlong Cheng · 2018
Later among the works it cites.
Yang You, Zhao Zhang, Cho-Jui Hsieh, and James Demmel · 2018
Later among the works it cites.
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2018
Later among the works it cites.
On the relation between the sharpest directions of DNN loss and the SGD step length
Stanisław Jastrzębski, Zachary Kenton, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amost Storkey · 2019
Closest in time.