Fetching the paper…
Reading the bibliography…
With the increasing demand for large-scale training of machine learning models, consensus-based distributed optimization methods have recently been advocated as alternatives to the popular parameter server framework.
Distributed Asynchronous Deterministic and Stochastic Gradient Optimization Algorithms
John Tsitsiklis, Dimitri Bertsekas, and Michael Athans · 1986
Earlier work this paper cites.
Principal Component Analysis
Svante Wold, Kim Esbensen, and Paul Geladi · 1987
Earlier work this paper cites.
Parallel and Distributed Computation: Numerical Methods
Dimitri P Bertsekas and John N Tsitsiklis · 1989
Earlier work this paper cites.
A Bridging Model for Parallel Computation
Leslie G Valiant · 1990
Earlier work this paper cites.
Gradient-Based Learning Applied to Document Recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Gossip-Based Computation of Aggregate Information
David Kempe, Alin Dobra, and Johannes Gehrke · 2003
Earlier work this paper cites.
Fast Linear Iterations for Distributed Averaging
Lin Xiao and Stephen Boyd · 2004
Earlier work this paper cites.
Randomized Gossip Algorithms
Stephen Boyd, Arpita Ghosh, Balaji Prabhakar, and Devavrat Shah · 2006
Earlier work this paper cites.
Distributed Average Consensus with Time-Varying Metropolis Weights
Lin Xiao, Stephen Boyd, and Sanjay Lall · 2006
Earlier work this paper cites.
Distributed Subgradient Methods for Multi-Agent Optimization
Angelia Nedic and Asuman Ozdaglar · 2009
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
An Architecture for Parallel Topic Models
Alexander Smola and Shravan Narayanamurthy · 2010
Earlier work this paper cites.
A Randomized Incremental Subgradient Method for Distributed Optimization in Networked Systems
Björn Johansson, Maben Rabi, and Mikael Johansson · 2010
Earlier work this paper cites.
Distributed Stochastic Subgradient Projection Algorithms for Convex Optimization
S Sundhar Ram, Angelia Nedić, and Venugopal V Veeravalli · 2010
Earlier work this paper cites.
Scaling Up Machine Learning: Parallel and Distributed Approaches
Ron Bekkerman, Mikhail Bilenko, and John Langford · 2011
Earlier work this paper cites.
Distributed Optimization and Statistical Learning via the Alternating Direction Method of Multipliers
Stephen Boyd, Neal Parikh, and Eric Chu · 2011
Earlier work this paper cites.
Dual Averaging for Distributed Optimization: Convergence Analysis and Network Scaling
John C Duchi, Alekh Agarwal, and Martin J Wainwright · 2011
Earlier work this paper cites.
Large Scale Distributed Deep Networks
Jeffrey Dean, Greg S Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V Le, Mark Z Mao, Marc’Aurelio Ranzato, Andrew Senior, Paul Tucker, et al · 2012
Cited alongside, same era.
More Effective Distributed ML Via A Stale Synchronous Parallel Parameter Server
Qirong Ho, James Cipar, Henggang Cui, Jin Kyu Kim, Seunghak Lee, Phillip B Gibbons, Garth A Gibson, Gregory R Ganger, and Eric P Xing · 2013
Cited alongside, same era.
Effective Straggler Mitigation: Attack of the Clones
Ganesh Ananthanarayanan, Ali Ghodsi, Scott Shenker, and Ion Stoica · 2013
Cited alongside, same era.
The Tail at Scale
Jeffrey Dean and Luiz André Barroso · 2013
Cited alongside, same era.
Scaling Distributed Machine Learning with The Parameter Server
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su · 2014
Cited alongside, same era.
Communication Efficient Distributed Machine Learning with the Parameter Server
Optimal Algorithms for Smooth and Strongly Convex Distributed Optimization in Networks
Kevin Scaman, Francis Bach, Sébastien Bubeck, Yin Tat Lee, and Laurent Massoulié · 2017
Later among the works it cites.
Train Longer, Generalize Better: Closing the Generalization Gap in Large Batch Training of Neural Networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Later among the works it cites.
Asynchronous Decentralized Parallel Stochastic Gradient Descent
Xiangru Lian, Wei Zhang, Ce Zhang, and Ji Liu · 2018
Later among the works it cites.
Optimal Algorithms for Non-Smooth Distributed Optimization in Networks
Kevin Scaman, Francis Bach, Sébastien Bubeck, Laurent Massoulié, and Yin Tat Lee · 2018
Later among the works it cites.
Communication Compression for Decentralized Training
Hanlin Tang, Shaoduo Gan, Ce Zhang, Tong Zhang, and Ji Liu · 2018
Later among the works it cites.
Decentralized Training Over Decentralized Data
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mu Li, David G Andersen, Alexander J Smola, and Kai Yu · 2014
Cited alongside, same era.
Convex Optimization: Algorithms and Complexity
Sébastien Bubeck et al · 2015
Cited alongside, same era.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Cited alongside, same era.
Gossip Dual Averaging for Decentralized Optimization of Pairwise Functions
Igor Colin, Aurélien Bellet, Joseph Salmon, and Stéphan Clémençon · 2016
Cited alongside, same era.
Revisiting Distributed Synchronous SGD
Jianmin Chen, Xinghao Pan, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Cited alongside, same era.
Tensorflow: A System for Large-Scale Machine Learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Hanlin Tang, Xiangru Lian, Ming Yan, Ce Zhang, and Ji Liu · 2018
Later among the works it cites.
Toward Understanding the Impact of Staleness in Distributed Machine Learning
Wei Dai, Yi Zhou, Nanqing Dong, Hao Zhang, and Eric Xing · 2018
Later among the works it cites.
Bayesian Distributed Stochastic Gradient Descent
Michael Teng and Frank Wood · 2018
Later among the works it cites.
Network Topology and Communication-Computation Tradeoffs in Decentralized Optimization
Angelia Nedić, Alex Olshevsky, and Michael G Rabbat · 2018
Later among the works it cites.
Optimization Methods for Large-Scale Machine Learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Later among the works it cites.
Don’t Decay the Learning Rate, Increase the Batch Size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2018
Later among the works it cites.
Hop: Heterogeneity-Aware Decentralized Training
Qinyi Luo, Jinkun Lin, Youwei Zhuo, and Xuehai Qian · 2019
Later among the works it cites.
Advances and Open Problems in Federated Learning
Peter Kairouz, H Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al · 2019
Later among the works it cites.
Towards Flexible Device Participation in Federated Learning for Non-IID Data
Yichen Ruan, Xiaoxi Zhang, Shu-Che Liang, and Carlee Joe-Wong · 2020
Later among the works it cites.
Dynamic Backup Workers for Parallel Machine Learning
Chuan Xu, Giovanni Neglia, and Nicola Sebastianelli · 2020
Later among the works it cites.
Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
Zhenheng Tang, Shaohuai Shi, Xiaowen Chu, Wei Wang, and Bo Li · 2020
Later among the works it cites.
Achieving Linear Speedup with Partial Worker Participation in Non-IID Federated Learning
Haibo Yang, Minghong Fang, and Jia Liu · 2021
Closest in time.