Fetching the paper…
Reading the bibliography…
We propose Stochastic Weight Averaging in Parallel (SWAP), an algorithm to accelerate DNN training.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Parallel sgd: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Earlier work this paper cites.
Adabatch: Adaptive batch sizes for training deep neural networks
Aditya Devarakonda, Maxim Naumov, and Michael Garland · 2017
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout
Terrance DeVries and Graham W Taylor · 2017
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2017
Earlier work this paper cites.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Cited alongside, same era.
Scaling sgd batch size to 32k for imagenet training
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Cited alongside, same era.
On the computational inefficiency of large batch sizes for stochastic gradient descent
Noah Golmant, Nikita Vemuri, Zhewei Yao, Vladimir Feinberg, Amir Gholami, Kai Rothauge, Michael W Mahoney, and Joseph Gonzalez · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Horovod: fast and easy distributed deep learning in tensorflow
Alexander Sergeev and Mike Del Balso · 2018
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
Christopher J Shallue, Jaehoon Lee, Joe Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl · 2018
Later among the works it cites.
Local sgd converges fast and communicates little
Sebastian U Stich · 2018
Later among the works it cites.
There are many consistent explanations of unlabeled data: Why you should average
Ben Athiwaratkun, Marc Finzi, Pavel Izmailov, and Andrew Gordon Wilson · 2019
Later among the works it cites.
Stochastic gradient methods with layer-wise adaptive moments for training of deep networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Xianyan Jia, Shutao Song, Wei He, Yangzihao Wang, Haidong Rong, Feihu Zhou, Liqiang Xie, Zhenyu Guo, Yuanzhou Yang, Liwei Yu, et al · 2018
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Cited alongside, same era.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U Stich, Kumar Kshitij Patel, and Martin Jaggi · 2018
Cited alongside, same era.
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Cited alongside, same era.
Dawnbench: An end-to-end deep learning benchmark and competition
Cody Coleman, Deepak Narayanan, Daniel Kang, Tian Zhao, Jian Zhang, Luigi Nardi, Peter Bailis, Kunle Olukotun, Chris Ré, and Matei Zaharia
Cited in the paper.
Improving stability in deep reinforcement learning with weight averaging
Evgenii Nikishin, Pavel Izmailov, Ben Athiwaratkun, Dmitrii Podoprikhin, Timur Garipov, Pavel Shvechikov, Dmitry Vetrov, and Andrew Gordon Wilson
Cited in the paper.
Boris Ginsburg, Patrice Castonguay, Oleksii Hrinchuk, Oleksii Kuchaiev, Vitaly Lavrukhin, Ryan Leary, Jason Li, Huyen Nguyen, and Jonathan M Cohen · 2019
Later among the works it cites.
Federated learning: Challenges, methods, and future directions
Tian Li, Anit Kumar Sahu, Ameet Talwalkar, and Virginia Smith · 2019
Later among the works it cites.
A simple baseline for bayesian uncertainty in deep learning
Wesley Maddox, Timur Garipov, Pavel Izmailov, Dmitry Vetrov, and Andrew Gordon Wilson · 2019
Later among the works it cites.
SWALP : Stochastic weight averaging in low precision training
Guandao Yang, Tianyi Zhang, Polina Kirichenko, Junwen Bai, Andrew Gordon Wilson, and Chris De Sa · 2019
Later among the works it cites.
Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning
Hao Yu, Sen Yang, and Shenghuo Zhu · 2019
Later among the works it cites.