Fetching the paper…
Reading the bibliography…
Decentralized algorithm is a form of computation that achieves a global goal through local dynamics that relies on low-cost communication between directly-connected agents.
“Language models are few-shot learners,”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei, · 1901
Earlier work this paper cites.
“Learning internal representations by error propagation,”
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams, · 1985
Earlier work this paper cites.
“Distributed asynchronous deterministic and stochastic gradient optimization algorithms,”
John Tsitsiklis, Dimitri Bertsekas, and Michael Athans, · 1986
Earlier work this paper cites.
“Backpropagation through time: what it does and how to do it,”
Paul J Werbos, · 1990
Earlier work this paper cites.
Using MPI-2: advanced features of the message passing interface
William Gropp, Rajeev Thakur, and Ewing Lusk, · 1999
Earlier work this paper cites.
“Decentralized optimization, with application to multiple aircraft coordination,”
Gokhan Inalhan, Dusan M Stipanovic, and Claire J Tomlin, · 2002
Earlier work this paper cites.
“Open MPI: Goals, concept, and design of a next generation MPI implementation,”
Edgar Gabriel, Graham E. Fagg, George Bosilca, Thara Angskun, Jack J. Dongarra, Jeffrey M. Squyres, Vishal Sahay, Prabhanjan Kambadur, Brian Barrett, Andrew Lumsdaine, Ralph H. Castain, David J. Daniel, Richard L. Graham, and Timothy S. Woodall, · 2004
Earlier work this paper cites.
“Complex networks and decentralized search algorithms,”
Jon Kleinberg, · 2006
Earlier work this paper cites.
“Diffusion least-mean squares over adaptive networks: Formulation and performance analysis,”
Cassio G Lopes and Ali H Sayed, · 2008
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database,”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, · 2009
Earlier work this paper cites.
“Distributed subgradient methods for multi-agent optimization,”
Angelia Nedic and Asuman Ozdaglar, · 2009
Earlier work this paper cites.
“Bandwidth optimal all-reduce algorithms for clusters of workstations,”
Pitch Patarasuk and Xin Yuan, · 2009
Earlier work this paper cites.
“An architecture for parallel topic models,”
Alexander Smola and Shravan Narayanamurthy, · 2010
Earlier work this paper cites.
“Distributed sparse linear regression,”
Gonzalo Mateos, Juan Andrés Bazerque, and Georgios B Giannakis, · 2010
Earlier work this paper cites.
“Weighted gossip: Distributed averaging using non-doubly stochastic matrices,”
Florence Bénézit, Vincent Blondel, Patrick Thiran, John Tsitsiklis, and Martin Vetterli, · 2010
Earlier work this paper cites.
“Bayesian learning in social networks,”
Daron Acemoglu, Munther A Dahleh, Ilan Lobel, and Asuman Ozdaglar, · 2011
Earlier work this paper cites.
“Centralized and decentralized control for demand response,”
Shuai Lu, Nader Samaan, Ruisheng Diao, Marcelo Elizondo, Chunlian Jin, Ebony Mayhorn, Yu Zhang, and Harold Kirkham, · 2011
Earlier work this paper cites.
“Dual averaging for distributed optimization: Convergence analysis and network scaling,”
John C Duchi, Alekh Agarwal, and Martin J Wainwright, · 2011
Earlier work this paper cites.
“Cooperative prey herding based on diffusion adaptation,”
Sheng-Yuan Tu and Ali H Sayed, · 2011
Earlier work this paper cites.
“Diffusion adaptation strategies for distributed optimization and learning over networks,”
Jianshu Chen and Ali H Sayed, · 2012
Earlier work this paper cites.
C++ concurrency in action
Anthony Williams, · 2012
Earlier work this paper cites.
“MPI: A Message-Passing Interface Standard Version 3.0,” 09 2012,
Message Passing Interface Forum, · 2012
Earlier work this paper cites.
“Complexity of Word Collocation Networks: A Preliminary Structural Analysis,”
Shibamouli Lahiri, · 2014
Earlier work this paper cites.
“Adaptation, learning, and optimization over networks,”
Ali H Sayed, · 2014
Earlier work this paper cites.
“On the linear convergence of the admm in decentralized consensus optimization,”
Wei Shi, Qing Ling, Kun Yuan, Gang Wu, and Wotao Yin, · 2014
Earlier work this paper cites.
“Scaling distributed machine learning with the parameter server,”
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su, · 2014
Earlier work this paper cites.
Using advanced MPI: Modern features of the message-passing interface
William Gropp, Torsten Hoefler, Rajeev Thakur, and Ewing Lusk, · 2014
Earlier work this paper cites.
“Distributed optimization over time-varying directed graphs,”
Angelia Nedić and Alex Olshevsky, · 2014
Earlier work this paper cites.
“Very deep convolutional networks for large-scale image recognition,”
Karen Simonyan and Andrew Zisserman, · 2014
Earlier work this paper cites.
“Extra: An exact first-order algorithm for decentralized consensus optimization,”
Wei Shi, Qing Ling, Gang Wu, and Wotao Yin, · 2015
Cited alongside, same era.
“A large contextual dataset for classification, detection and counting of cars with deep learning,”
T Nathan Mundhenk, Goran Konjevod, Wesam A Sakla, and Kofi Boakye, · 2016
Cited alongside, same era.
“Next: In-network nonconvex optimization,”
P. Di Lorenzo and G. Scutari, · 2016
Cited alongside, same era.
“Tensorflow: A system for large-scale machine learning,”
Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al., · 2016
Cited alongside, same era.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Cited alongside, same era.
“Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent,”
“Doublesqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression,”
Hanlin Tang, Chen Yu, Xiangru Lian, Tong Zhang, and Ji Liu, · 2019
Later among the works it cites.
“Local sgd converges fast and communicates little,”
Sebastian Urban Stich, · 2019
Later among the works it cites.
“On the linear speedup analysis of communication efficient momentum sgd for distributed non-convex optimization,”
Hao Yu, Rong Jin, and Sen Yang, · 2019
Later among the works it cites.
“Communication-censored admm for decentralized consensus optimization,”
Yaohua Liu, Wei Xu, Gang Wu, Zhi Tian, and Qing Ling, · 2019
Later among the works it cites.
“Pytorch: An imperative style, high-performance deep learning library,”
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al., · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu, · 2017
Cited alongside, same era.
“Achieving geometric convergence for distributed optimization over time-varying graphs,”
Angelia Nedic, Alex Olshevsky, and Wei Shi, · 2017
Cited alongside, same era.
“Optimal algorithms for smooth and strongly convex distributed optimization in networks,”
Kevin Scaman, Francis Bach, Sébastien Bubeck, Yin Tat Lee, and Laurent Massoulié, · 2017
Cited alongside, same era.
“Qsgd: Communication-efficient sgd via gradient quantization and encoding,”
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic, · 2017
Cited alongside, same era.
“Accurate, large minibatch sgd: Training imagenet in 1 hour,”
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He, · 2017
Cited alongside, same era.
“Asynchronous decentralized parallel stochastic gradient descent,”
Xiangru Lian, Wei Zhang, Ce Zhang, and Ji Liu, · 2018
Cited alongside, same era.
“Exact diffusion for distributed optimization and learning—part i: Algorithm development,”
Kun Yuan, Bicheng Ying, Xiaochuan Zhao, and Ali H Sayed, · 2018
Cited alongside, same era.
“Efficientnet: Rethinking model scaling for convolutional neural networks,”
Mingxing Tan and Quoc Le, · 2019
Later among the works it cites.
“Bringing hpc techniques to deep learning,” https://andrew.gibiansky.com/blog/machine-learning/baidu-allreduce/ , 2017,
Andrew Gibiansky, · 2020
Later among the works it cites.
“Decentralized stochastic non-convex optimization over weakly connected time-varying digraphs,”
Songtao Lu and Chai Wah Wu, · 2020
Later among the works it cites.
“A unified theory of decentralized sgd with changing topology and local updates,”
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian U Stich, · 2020
Later among the works it cites.
“A unified architecture for accelerating distributed {DNN} training in heterogeneous gpu/cpu clusters,”
Yimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi, Yong Cui, and Chuanxiong Guo, · 2020
Later among the works it cites.
“A dual approach for optimal algorithms in distributed optimization over networks,”
César A Uribe, Soomin Lee, Alexander Gasnikov, and Angelia Nedić, · 2020
Later among the works it cites.
“An improved convergence analysis for decentralized online stochastic non-convex optimization,”
Ran Xin, Usman A Khan, and Soummya Kar, · 2020
Later among the works it cites.
“Prague: High-performance heterogeneity-aware asynchronous decentralized training,”
Qinyi Luo, Jiaao He, Youwei Zhuo, and Xuehai Qian, · 2020
Later among the works it cites.
“Push–pull gradient methods for distributed optimization in networks,”
Shi Pu, Wei Shi, Jinming Xu, and Angelia Nedić, · 2020
Later among the works it cites.
“Accelerating gossip SGD with periodic global averaging,”
Yiming Chen, Kun Yuan, Yingya Zhang, Pan Pan, Yinghui Xu, and Wotao Yin, · 2021
Closest in time.
“DecentLaM: Decentralized momentum SGD for large-batch deep training,”
Kun Yuan, Yiming Chen, Xinmeng Huang, Yingya Zhang, Pan Pan, Yinghui Xu, and Wotao Yin, · 2021
Closest in time.
“Consensus control for decentralized deep learning,”
Lingjing Kong, Tao Lin, Anastasia Koloskova, Martin Jaggi, and Sebastian U Stich, · 2021
Closest in time.
“A unified and refined convergence analysis for non-convex decentralized learning,”
Sulaiman A Alghunaim and Kun Yuan, · 2021
Closest in time.
Large-Scale Convex Optimization via Monotone Operators
Ernest K. Ryu and Wotao Yin, · 2021
Closest in time.
“Exponential graph is provably efficient for decentralized deep training,”
Bicheng Ying, Kun Yuan, Yiming Chen, Hanbin Hu, Pan Pan, and Wotao Yin, · 2021
Closest in time.
“Improving the transient times for distributed stochastic gradient methods,”
Kun Huang and Shi Pu, · 2021
Closest in time.
“Removing data heterogeneity influence enhances network topology dependence of decentralized sgd,”
Kun Yuan and Sulaiman A Alghunaim, · 2021
Closest in time.
“Relaysum for decentralized deep learning on heterogeneous data,”
Thijs Vogels, Lie He, Anastasia Koloskova, Tao Lin, Sai Praneeth Karimireddy, Sebastian U Stich, and Martin Jaggi, · 2021
Closest in time.
“Optimal complexity in decentralized training,”
Yucheng Lu and Christopher De Sa, · 2021
Closest in time.
Ran Xin, Subhro Das, Usman A Khan, and Soummya Kar, · 2021
Closest in time.
“Quasi-global momentum: Accelerating decentralized deep learning on heterogeneous data,”
Tao Lin, Sai Praneeth Karimireddy, Sebastian U Stich, and Martin Jaggi, · 2021
Closest in time.
“Bagua: Scaling up distributed learning with system relaxations,”
Shaoduo Gan, Xiangru Lian, Rui Wang, Jianbin Chang, Chengjun Liu, Hongmei Shi, Shengzhuo Zhang, Xianghong Li, Tengxu Sun, Jiawei Jiang, et al., · 2021
Closest in time.
“Augmented distributed gradient methods for multi-agent optimization under uncoordinated constant stepsizes,”
J. Xu, S. Zhu, Y. C. Soh, and L. Xie, · 2060
Closest in time.