Fetching the paper…
Reading the bibliography…
1-bit gradient compression and local steps are two representative techniques that enable drastic communication reduction in distributed SGD.
Sadam: A variant of adam for strongly convex functions
Guanghui Wang, Shiyin Lu, Weiwei Tu, and Lijun Zhang · 1905
Earlier work this paper cites.
Distributed training with heterogeneous data: Bridging median-and mean-based algorithms
Xiangyi Chen, Tiancong Chen, Haoran Sun, Zhiwei Steven Wu, and Mingyi Hong · 1906
Earlier work this paper cites.
Zo-adamm: Zeroth-order adaptive momentum method for black-box optimization
Xiangyi Chen, Sijia Liu, Kaidi Xu, Xingguo Li, Xue Lin, Mingyi Hong, and David Cox · 1910
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Feng Niu, Benjamin Recht, Christopher Ré, and Stephen J Wright · 2011
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Torchvision 0.11.0 documentation — pytorch.org
Pytorch · 2014
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Earlier work this paper cites.
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Earlier work this paper cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Earlier work this paper cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Local sgd converges fast and communicates little
Sebastian U Stich · 2018
Earlier work this paper cites.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U Stich, Kumar Kshitij Patel, and Martin Jaggi · 2018
Earlier work this paper cites.
signsgd via zeroth-order oracle
Sijia Liu, Pin-Yu Chen, Xiangyi Chen, and Mingyi Hong · 2018
Cited alongside, same era.
Dissecting adam: The sign, magnitude and variance of stochastic gradients
Lukas Balles and Philipp Hennig · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2018
Cited alongside, same era.
Nostalgic adam: Weighting more of the past gradients when designing the adaptive learning rate
Haiwen Huang, Chang Wang, and Bin Dong · 2018
Cited alongside, same era.
Distributed signsgd with improved accuracy and network-fault tolerance
Trieu Le Phong and Tran Thi Phuong · 2020
Later among the works it cites.
Stochastic-sign sgd for federated learning with theoretical guarantees
Richeng Jin, Yufan Huang, Xiaofan He, Huaiyu Dai, and Tianfu Wu · 2020
Later among the works it cites.
Moniqua: Modulo quantized communication in decentralized sgd
Yucheng Lu and Christopher De Sa · 2020
Later among the works it cites.
A new regret analysis for adam-type algorithms
Ahmet Alacaoglu, Yura Malitsky, Panayotis Mertikopoulos, and Volkan Cevher · 2020
Later among the works it cites.
A simple convergence proof of adam and adagrad
Alexandre Défossez, Léon Bottou, Francis Bach, and Nicolas Usunier · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Trieu H Trinh and Quoc V Le · 2018
Cited alongside, same era.
signadam++: Learning confidences for deep neural networks
Dong Wang, Yicheng Liu, Wenwo Tang, Fanhua Shang, Hongying Liu, Qigong Sun, and Licheng Jiao · 2019
Cited alongside, same era.
Signprox: One-bit proximal algorithm for nonconvex stochastic optimization
Xiaojian Xu and Ulugbek S Kamilov · 2019
Cited alongside, same era.
Error feedback fixes signsgd and other gradient compression schemes
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi · 2019
Cited alongside, same era.
Election coding for distributed learning: Protecting signsgd against byzantine attacks
Jy-yong Sohn, Dong-Jun Han, Beongjun Choi, and Jaekyun Moon · 2019
Cited alongside, same era.
Decentralized deep learning with arbitrary communication compression
Anastasia Koloskova, Tao Lin, Sebastian U Stich, and Martin Jaggi · 2019
Cited alongside, same era.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2019
Cited alongside, same era.
Yucheng Lu, Jack Nash, and Christopher De Sa · 2020
Later among the works it cites.
Recent theoretical advances in non-convex optimization
Marina Danilova, Pavel Dvurechensky, Alexander Gasnikov, Eduard Gorbunov, Sergey Guminov, Dmitry Kamzolov, and Innokentiy Shibaev · 2020
Later among the works it cites.
Adabelief optimizer: Adapting stepsizes by the belief in observed gradients
Juntang Zhuang, Tommy Tang, Yifan Ding, Sekhar C Tatikonda, Nicha Dvornek, Xenophon Papademetris, and James Duncan · 2020
Later among the works it cites.
Qsparse-local-sgd: Distributed sgd with quantization, sparsification, and local computations
Debraj Basu, Deepesh Data, Can Karakus, and Suhas N Diggavi · 2020
Later among the works it cites.
1-bit adam: Communication efficient large-scale training with adam’s convergence speed
Hanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari, Conglong Li, Xiangru Lian, Ji Liu, Ce Zhang, and Yuxiong He · 2021
Later among the works it cites.
Optimal complexity in decentralized training
Yucheng Lu and Christopher De Sa · 2021
Later among the works it cites.
Stochastic sign descent methods: New algorithms and better theory
Mher Safaryan and Peter Richtárik · 2021
Later among the works it cites.
Dp-signsgd: When efficiency meets privacy and robustness
Lingjuan Lyu · 2021
Later among the works it cites.
Federated learning via plurality vote
Kai Yue, Richeng Jin, Chau-Wai Wong, and Huaiyu Dai · 2021
Later among the works it cites.
Sign-maml: Efficient model-agnostic meta-learning by signsgd
Chen Fan, Parikshit Ram, and Sijia Liu · 2021
Later among the works it cites.
Momentum centering and asynchronous update for adaptive gradient methods
Juntang Zhuang, Yifan Ding, Tommy Tang, Nicha Dvornek, Sekhar C Tatikonda, and James Duncan · 2021
Later among the works it cites.
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zheng, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michael Houston, Saurabh Tiwary, and Bryan Catanzaro · 2022
Closest in time.