Fetching the paper…
Reading the bibliography…
In decentralized machine learning, workers compute model updates on their local data.
Deep South: A social anthropological study of caste and class
Allison Davis, Burleigh Bradford Gardner, and Mary R Gardner · 1930
Earlier work this paper cites.
An algorithm for distributed computation of a spanningtree in an extended LAN
Radia J. Perlman · 1985
Earlier work this paper cites.
Gossip-based computation of aggregate information
David Kempe, Alin Dobra, and Johannes Gehrke · 2003
Earlier work this paper cites.
Fast linear iterations for distributed averaging
Lin Xiao and Stephen P. Boyd · 2004
Earlier work this paper cites.
Exploring network structure, dynamics, and function using NetworkX
Aric Hagberg, Pieter Swart, and Daniel S Chult · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li · 2009
Earlier work this paper cites.
A randomized incremental subgradient method for distributed optimization in networked systems
Björn Johansson, Maben Rabi, and Mikael Johansson · 2009
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Angelia Nedic and Asuman E. Ozdaglar · 2009
Earlier work this paper cites.
Two-tree algorithms for full bandwidth broadcast, reduction and scan
Peter Sanders, Jochen Speck, and Jesper Larsson Träff · 2009
Earlier work this paper cites.
On Matrix Factorization and Scheduling forFinite-time Average-consensus
Chih-Kai Ko · 2010
Earlier work this paper cites.
Convergence rate for consensus with delays
Angelia Nedić and Asuman Ozdaglar · 2010
Earlier work this paper cites.
Definitive Consensus for Distributed Data Inference
Leonidas Georgopoulos · 2011
Earlier work this paper cites.
Distributed consensus and optimization under communication delays
Konstantinos I. Tsianos and Michael G. Rabbat · 2011
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C. Duchi · 2012
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2012
Earlier work this paper cites.
Push-sum distributed dual averaging for convex optimization
Konstantinos I. Tsianos, Sean F. Lawlor, and Michael G. Rabbat · 2012
Earlier work this paper cites.
Graph diameter, eigenvalues, and minimum-time consensus
Julien M. Hendrickx, Raphaël M. Jungers, Alexander Olshevsky, and Guillaume Vankeerberghen · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Jake Zhao, and Yann LeCun · 2015
Cited alongside, same era.
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
Next: In-network nonconvex optimization
Paolo Di Lorenzo and Gesualdo Scutari · 2016
Cited alongside, same era.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Later among the works it cites.
Unified optimal analysis of the (stochastic) gradient method
Sebastian U. Stich · 2019
Later among the works it cites.
FROST - fast row-stochastic optimization with uncoordinated step-sizes
Ran Xin, Chenguang Xi, and Usman A. Khan · 2019
Later among the works it cites.
Exact diffusion for distributed optimization and learning - part I: algorithm development
Kun Yuan, Bicheng Ying, Xiaochuan Zhao, and Ali H. Sayed · 2019
Later among the works it cites.
Bayesian nonparametric federated learning of neural networks
Mikhail Yurochkin, Mayank Agarwal, Soumya Ghosh, Kristjan H. Greenewald, Trong Nghia Hoang, and Yasaman Khazaeni · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic gradient-push for strongly convex functions on time-varying directed graphs
Angelia Nedic and Alex Olshevsky · 2016
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? A case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas · 2017
Cited alongside, same era.
Achieving geometric convergence for distributed optimization over time-varying graphs
Angelia Nedic, Alex Olshevsky, and Wei Shi · 2017
Cited alongside, same era.
DEXTRA: A fast algorithm for optimization over directed graphs
Chenguang Xi and Usman A. Khan · 2017
Cited alongside, same era.
Distributed stochastic gradient tracking methods
Shi Pu and Angelia Nedic · 2018
Cited alongside, same era.
A bp-like distributed algorithm for weighted average consensus
Zhaorong Zhang, Kan Xie, Qianqian Cai, and Minyue Fu · 2019
Later among the works it cites.
A unified theory of decentralized SGD with changing topology and local updates
Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian U. Stich · 2020
Later among the works it cites.
Evolving normalization-activation layers
Hanxiao Liu, Andy Brock, Karen Simonyan, and Quoc Le · 2020
Later among the works it cites.
Distributed gradient methods for convex machine learning problems in networks: Distributed optimization
Angelia Nedic · 2020
Later among the works it cites.
Distributed heavy-ball: A generalization and acceleration of first-order methods with gradient tracking
Ran Xin and Usman A. Khan · 2020
Later among the works it cites.
Decentralized stochastic gradient tracking for non-convex empirical risk minimization, 2020
Jiaqi Zhang and Keyou You · 2020
Later among the works it cites.
Quasi-global momentum: Accelerating decentralized deep learning on heterogeneous data
Tao Lin, Sai Praneeth Karimireddy, Sebastian U. Stich, and Martin Jaggi · 2021
Closest in time.
Optimal complexity in decentralized training
Yucheng Lu and Christopher De Sa · 2021
Closest in time.
Push-pull gradient methods for distributed optimization in networks
Shi Pu, Wei Shi, Jinming Xu, and Angelia Nedic · 2021
Closest in time.
Decentlam: Decentralized momentum SGD for large-batch deep training
Kun Yuan, Yiming Chen, Xinmeng Huang, Yingya Zhang, Pan Pan, Yinghui Xu, and Wotao Yin · 2021
Closest in time.