Fetching the paper…
Reading the bibliography…
Decentralized optimization is emerging as a viable alternative for scalable distributed machine learning, but also introduces new challenges in terms of synchronization costs.
Decentralized deep learning with arbitrary communication compression
Anastasia Koloskova, Tao Lin, Sebastian U Stich, and Martin Jaggi · 1907
Earlier work this paper cites.
Decentralized deep learning with arbitrary communication compression
Anastasia Koloskova, Tao Lin, Sebastian U Stich, and Martin Jaggi · 1907
Earlier work this paper cites.
Problems in decentralized decision making and computation
John Nikolas Tsitsiklis · 1984
Earlier work this paper cites.
Dynamic load balancing by random matchings
Bhaskar Ghosh and S. Muthukrishnan · 1996
Earlier work this paper cites.
Simple, fast, and practical non-blocking and blocking concurrent queue algorithms
Maged M. Michael and Michael L. Scott · 1996
Earlier work this paper cites.
Gossip-based computation of aggregate information
David Kempe, Alin Dobra, and Johannes Gehrke · 2003
Earlier work this paper cites.
Fast linear iterations for distributed averaging
Lin Xiao and Stephen Boyd · 2004
Earlier work this paper cites.
Computation in networks of passively mobile finite-state sensors
Dana Angluin, James Aspnes, Zoë Diamadi, Michael J Fischer, and René Peralta · 2006
Earlier work this paper cites.
Randomized gossip algorithms
Stephen Boyd, Arpita Ghosh, Balaji Prabhakar, and Devavrat Shah · 2006
Earlier work this paper cites.
High performance rdma protocols in hpc
Tim S Woodall, Galen M Shipman, George Bosilca, Richard L Graham, and Arthur B Maccabe · 2006
Earlier work this paper cites.
A new analytical method for parallel, diffusion-type load balancing
Petra Berenbrink, Tom Friedetzky, and Zengjian Hu · 2008
Earlier work this paper cites.
Technology-driven, highly-scalable dragonfly topology
John Kim, Wiliam J Dally, Steve Scott, and Dennis Abts · 2008
Earlier work this paper cites.
A randomized incremental subgradient method for distributed optimization in networked systems
Björn Johansson, Maben Rabi, and Mikael Johansson · 2009
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Angelia Nedic and Asuman Ozdaglar · 2009
Cited alongside, same era.
Slim fly: A cost effective low-diameter network topology
Maciej Besta and Torsten Hoefler · 2014
Cited alongside, same era.
The cifar-10 dataset
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton · 2014
Cited alongside, same era.
Distributed stochastic optimization and learning
Ohad Shamir and Nathan Srebro · 2014
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Cited alongside, same era.
Asynchronous decentralized parallel stochastic gradient descent
Xiangru Lian, Wei Zhang, Ce Zhang, and Ji Liu · 2018
Later among the works it cites.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U Stich, Kumar Kshitij Patel, and Martin Jaggi · 2018
Later among the works it cites.
Local sgd converges fast and communicates little
Sebastian U Stich · 2018
Later among the works it cites.
Decentralization meets quantization
Hanlin Tang, Ce Zhang, Shaoduo Gan, Tong Zhang, and Ji Liu · 2018
Later among the works it cites.
Jianyu Wang and Gauri Joshi · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
QSGD: Randomized quantization for communication-efficient stochastic gradient descent
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jio Hsieh, Wei Zhang, and Ji Liu · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Stochastic gradient push for distributed deep learning
Mahmoud Assran, Nicolas Loizou, Nicolas Ballas, and Michael Rabbat · 2018
Cited alongside, same era.
Later among the works it cites.
Taming unbalanced training workloads in deep learning with partial collective operations
Shigang Li, Tal Ben-Nun, Salvatore Di Girolamo, Dan Alistarh, and Torsten Hoefler · 2019
Closest in time.
Slowmo: Improving communication-efficient distributed sgd with slow momentum
Jianyu Wang, Vinayak Tantia, Nicolas Ballas, and Michael Rabbat · 2019
Closest in time.
http://www.cscs.ch/computers/piz_daint , 2019
The CSCS Piz Daint supercomputer · 2020
Closest in time.
A unified theory of decentralized sgd with changing topology and local updates
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian U Stich · 2020
Closest in time.
Moniqua: Modulo quantized communication in decentralized SGD
Yucheng Lu and Christopher De Sa · 2020
Closest in time.
Distributed variance reduction with optimal communication
Peter Davies, Vijaykrishna Gurunathan, Niusha Moshrefi, Saleh Ashkboos Ashkboos, and Dan Alistarh · 2021
Closest in time.
Fully asynchronous distributed optimization with linear convergence in directed networks, 2021
Jiaqi Zhang and Keyou You · 2021
Closest in time.