Fetching the paper…
Reading the bibliography…
Distributed optimization often consists of two updating phases: local optimization and inter-node communication.
On rings of operators. reduction theory
John Von Neumann · 1949
Earlier work this paper cites.
Parallel and distributed computation: numerical methods
Dimitri P Bertsekas and John N Tsitsiklis · 1989
Earlier work this paper cites.
Block-iterative projection methods for parallel computation of solutions to convex feasibility problems
Ron Aharoni and Yair Censor · 1989
Earlier work this paper cites.
On projection algorithms for solving convex feasibility problems
Heinz H Bauschke and Jonathan M Borwein · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays
Uri Alon, Naama Barkai, Daniel A Notterman, Kurt Gish, Suzanne Ybarra, Daniel Mack, and Arnold J Levine · 1999
Earlier work this paper cites.
Convergence rate analysis and error bounds for projection algorithms in convex feasibility problems
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Dual averaging for distributed optimization: Convergence analysis and network scaling
John C Duchi, Alekh Agarwal, and Martin J Wainwright · 2012
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, Martin J Wainwright, and John C Duchi · 2012
Earlier work this paper cites.
More effective distributed ml via a stale synchronous parallel parameter server
Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B Gibbons, Garth A Gibson, Greg Ganger, and Eric P Xing · 2013
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su · 2014
Cited alongside, same era.
Communication efficient distributed machine learning with the parameter server
Mu Li, David G Andersen, Alexander J Smola, and Kai Yu · 2014
Cited alongside, same era.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Cited alongside, same era.
Communication quantization for data-parallel training of deep neural networks
Nikoli Dryden, Tim Moon, Sam Ade Jacobs, and Brian Van Essen · 2016
Cited alongside, same era.
Parallel sgd: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Cited alongside, same era.
Fan Zhou and Guojing Cong · 2017
Later among the works it cites.
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2017
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Later among the works it cites.
An optimal control approach to deep learning and applications to discrete-weight neural networks
Qianxiao Li and Shuji Hao · 2018
Later among the works it cites.
Cocoa: A general framework for communication-efficient distributed optimization
Virginia Smith, Simone Forte, Ma Chenxin, Martin Takac, Michael I Jordan, and Martin Jaggi · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
The implicit convex feasibility problem and its application to adaptive image denoising
Yair Censor, Aviv Gibali, Frank Lenzen, et al · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J Dally · 2017
Cited alongside, same era.
Distributed optimization with arbitrary local solvers
Chenxin Ma, Jakub Konecny, Martin Jaggi, Virginia Smith, Michael I Jordan, Peter Richtarik, and Martin Takac · 2017
Cited alongside, same era.
Later among the works it cites.
Local sgd converges fast and communicates little
Sebastian U Stich · 2018
Later among the works it cites.
Parallel restarted sgd for non-convex optimization with faster convergence and less communication
Hao Yu, Sen Yang, and Shenghuo Zhu · 2018
Later among the works it cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Later among the works it cites.
On exponential convergence of sgd in non-convex over-parametrized learning
Raef Bassily, Mikhail Belkin, and Siyuan Ma · 2018
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2018
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Later among the works it cites.