Fetching the paper…
Reading the bibliography…
The scale of deep learning nowadays calls for efficient distributed training algorithms.
Distributed asynchronous deterministic and stochastic gradient optimization algorithms
John Tsitsiklis, Dimitri Bertsekas, and Michael Athans · 1986
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Angelia Nedic and Asuman Ozdaglar · 2009
Earlier work this paper cites.
Bandwidth optimal all-reduce algorithms for clusters of workstations
Pitch Patarasuk and Xin Yuan · 2009
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman · 2010
Earlier work this paper cites.
Dual averaging for distributed optimization: Convergence analysis and network scaling
John C Duchi, Alekh Agarwal, and Martin J Wainwright · 2011
Earlier work this paper cites.
Diffusion adaptation strategies for distributed optimization and learning over networks
Jianshu Chen and Ali H Sayed · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Adaptation, learning, and optimization over networks
Ali H Sayed · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Faster r-cnn: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2016
Earlier work this paper cites.
On the convergence of decentralized gradient descent
Kun Yuan, Qing Ling, and Wotao Yin · 2016
Earlier work this paper cites.
On the influence of momentum acceleration on online learning
Kun Yuan, Bicheng Ying, and Ali H Sayed · 2016
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
signsgd with majority vote is communication efficient and fault tolerant
Jeremy Bernstein, Jiawei Zhao, Kamyar Azizzadenesheli, and Anima Anandkumar · 2018
Cited alongside, same era.
LAG: Lazily aggregated gradient for communication-efficient distributed learning
Tianyi Chen, Georgios Giannakis, Tao Sun, and Wotao Yin · 2018
Cited alongside, same era.
Autoaugment: Learning augmentation policies from data
Ekin D Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V Le · 2018
Cited alongside, same era.
Asynchronous decentralized parallel stochastic gradient descent
Xiangru Lian, Wei Zhang, Ce Zhang, and Ji Liu · 2018
Local sgd converges fast and communicates little
Sebastian Urban Stich · 2019
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Later among the works it cites.
Doublesqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression
Hanlin Tang, Chen Yu, Xiangru Lian, Tong Zhang, and Ji Liu · 2019
Later among the works it cites.
Slowmo: Improving communication-efficient distributed sgd with slow momentum
Jianyu Wang, Vinayak Tantia, Nicolas Ballas, and Michael Rabbat · 2019
Later among the works it cites.
Large-batch training for lstm and beyond
Yang You, Jonathan Hseu, Chris Ying, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2019
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Cited alongside, same era.
d 2 d^{2} : Decentralized training over decentralized data
Hanlin Tang, Xiangru Lian, Ming Yan, Ce Zhang, and Ji Liu · 2018
Cited alongside, same era.
Imagenet training in minutes
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer · 2018
Cited alongside, same era.
Exact diffusion for distributed optimization and learning—part ii: Convergence analysis
Kun Yuan, Bicheng Ying, Xiaochuan Zhao, and Ali H Sayed · 2018
Cited alongside, same era.
Stochastic gradient push for distributed deep learning
Mahmoud Assran, Nicolas Loizou, Nicolas Ballas, and Mike Rabbat · 2019
Cited alongside, same era.
Mmdetection: Open mmlab detection toolbox and benchmark
Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, et al · 2019
Cited alongside, same era.
Understanding the role of momentum in stochastic gradient methods
Igor Gitman, Hunter Lang, Pengchuan Zhang, and Lin Xiao · 2019
Cited alongside, same era.
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2019
Later among the works it cites.
On the linear speedup analysis of communication efficient momentum sgd for distributed non-convex optimization
Hao Yu, Rong Jin, and Sen Yang · 2019
Later among the works it cites.
Decentralized deep learning using momentum-accelerated consensus
Aditya Balu, Zhanhong Jiang, Sin Yong Tan, Chinmay Hedge, Young M Lee, and Soumik Sarkar · 2020
Later among the works it cites.
Periodic stochastic gradient descent with momentum for decentralized training
Hongchang Gao and Heng Huang · 2020
Later among the works it cites.
A unified theory of decentralized sgd with changing topology and local updates
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian U Stich · 2020
Later among the works it cites.
An improved analysis of stochastic gradient descent with momentum
Yanli Liu, Yuan Gao, and Wotao Yin · 2020
Later among the works it cites.
Momentum and stochastic momentum for stochastic gradient, newton, proximal point and subspace descent methods
Nicolas Loizou and Peter Richtárik · 2020
Later among the works it cites.
Prague: High-performance heterogeneity-aware asynchronous decentralized training
Qinyi Luo, Jiaao He, Youwei Zhuo, and Xuehai Qian · 2020
Later among the works it cites.
On the convergence of the stochastic heavy ball method
Othmane Sebbouh, Robert M Gower, and Aaron Defazio · 2020
Later among the works it cites.
Squarm-sgd: Communication-efficient momentum sgd for decentralized optimization
Navjot Singh, Deepesh Data, Jemin George, and Suhas Diggavi · 2020
Later among the works it cites.
An improved convergence analysis for decentralized online stochastic non-convex optimization
Ran Xin, Usman A Khan, and Soummya Kar · 2020
Later among the works it cites.
On the influence of bias-correction on distributed stochastic optimization
Kun Yuan, Sulaiman A Alghunaim, Bicheng Ying, and Ali H Sayed · 2020
Later among the works it cites.
Consensus control for decentralized deep learning
Lingjing Kong, Tao Lin, Anastasia Koloskova, Martin Jaggi, and Sebastian U Stich · 2021
Closest in time.
Quasi-global momentum: Accelerating decentralized deep learning on heterogeneous data
Tao Lin, Sai Praneeth Karimireddy, Sebastian U Stich, and Martin Jaggi · 2021
Closest in time.