Fetching the paper…
Reading the bibliography…
Distributed Stochastic Gradient Descent (SGD) when run in a synchronous manner, suffers from delays in runtime as it waits for the slowest workers (stragglers).
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Distributed asynchronous deterministic and stochastic gradient optimization algorithms
John Tsitsiklis, Dimitri Bertsekas, and Michael Athans · 1986
Earlier work this paper cites.
A course in microeconomic theory
David M Kreps · 1990
Earlier work this paper cites.
A first course in probability
Ross Sheldon · 2002
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean et al · 2012
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas L Roux, Mark Schmidt, and Francis R Bach · 2012
Earlier work this paper cites.
The tail at scale
Jeffrey Dean and Luiz Andre Barroso · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Solving the straggler problem with bounded staleness
James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Gregory R. Ganger, Garth Gibson, Kimberly Keeton, and Eric Xing · 2013
Earlier work this paper cites.
More effective distributed ml via a stale synchronous parallel parameter server
Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B Gibbons, Garth A Gibson, Greg Ganger, and Eric P Xing · 2013
Earlier work this paper cites.
Stochastic Processes: Theory for Applications
Robert G. Gallager · 2013
Earlier work this paper cites.
Efficient mini-batch training for stochastic optimization
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J. Smola · 2014
Earlier work this paper cites.
On the delay-storage trade-off in content download from coded distributed storage systems
Gauri Joshi, Yanpei Liu, and Emina Soljanin · 2014
Earlier work this paper cites.
Exploiting bounded staleness to speed up big data analytics
Henggang Cui, James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Abhimanu Kumar, Jinliang Wei, Wei Dai, Gregory R Ganger, Phillip B Gibbons, et al · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Using straggler replication to reduce latency in large-scale parallel computing
Da Wang, Gauri Joshi, and Gregory Wornell · 2015
Earlier work this paper cites.
Queues with redundancy: Latency-cost analysis
Gauri Joshi, Emina Soljanin, and Gregory Wornell · 2015
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Earlier work this paper cites.
Asynchronous stochastic convex optimization: the noise is in the noise and sgd don’t care
Sorathan Chaturapruek, John C Duchi, and Christopher Re · 2015
Earlier work this paper cites.
Deep learning with elastic averaging sgd
Sixin Zhang, Anna E Choromanska, and Yann LeCun · 2015
Cited alongside, same era.
Optimization methods for large-scale machine learning
Leon Bottou, Frank E. Curtis, and Jorge Nocedal · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Cited alongside, same era.
Addressing the straggler problem for iterative convergent parallel ml
Aaron Harlap, Henggang Cui, Wei Dai, Jinliang Wei, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, and Eric P. Xing · 2016
Cited alongside, same era.
Short-dot: Computing large linear transforms distributedly using coded short dot products
Sanghamitra Dutta, Viveck Cadambe, and Pulkit Grover · 2016
Cited alongside, same era.
Codes for Distributed Computing: A Tutorial
Viveck Cadambe and Pulkit Grover · 2017
Later among the works it cites.
Coded convolution for parallel and distributed computing within a deadline
Sanghamitra Dutta, Viveck Cadambe, and Pulkit Grover · 2017
Later among the works it cites.
Perturbed iterate analysis for asynchronous stochastic optimization
Horia Mania, Xinghao Pan, Dimitris Papailiopoulos, Benjamin Recht, Kannan Ramchandran, and Michael I Jordan · 2017
Later among the works it cites.
Robert Hannah and Wotao Yin · 2017
Later among the works it cites.
Asynchronous coordinate descent under more realistic assumptions
Tao Sun, Robert Hannah, and Wotao Yin · 2017
Later among the works it cites.
Asaga: Asynchronous parallel saga
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fault-tolerant distributed logistic regression using unreliable components
Yaoqing Yang, Pulkit Grover, and Soummya Kar · 2016
Cited alongside, same era.
Model accuracy and runtime tradeoff in distributed deep learning: A systematic study
Suyog Gupta, Wei Zhang, and Fei Wang · 2016
Cited alongside, same era.
Asynchrony begets momentum, with an application to deep learning
Ioannis Mitliagkas, Ce Zhang, Stefan Hadjis, and Christopher Re · 2016
Cited alongside, same era.
Staleness-aware async-sgd for distributed deep learning
Wei Zhang, Suyog Gupta, Xiangru Lian, and Ji Liu · 2016
Cited alongside, same era.
Arock: an algorithmic framework for asynchronous parallel coordinate updates
Zhimin Peng, Yangyang Xu, Ming Yan, and Wotao Yin · 2016
Cited alongside, same era.
On unbounded delays in asynchronous parallel fixed-point algorithms
Robert Hannah and Wotao Yin · 2016
Cited alongside, same era.
Revisiting distributed synchronous SGD
Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Cited alongside, same era.
Remi Leblond, Fabian Pedregosa, and Simon Lacoste-Julien · 2017
Later among the works it cites.
Gradient diversity empowers distributed learning
Dong Yin, Ashwin Pananjady, Max Lam, Dimitris Papailiopoulos, Kannan Ramchandran, and Peter Bartlett · 2017
Later among the works it cites.
Fan Zhou and Guojing Cong · 2017
Later among the works it cites.
Slow and stale gradients can win the race: Error-runtime trade-offs in distributed sgd
Sanghamitra Dutta, Gauri Joshi, Soumyadip Ghosh, Parijat Dube, and Priya Nagpurkar · 2018
Later among the works it cites.
Speeding up distributed machine learning using codes
Kangwook Lee, Maximilian Lam, Ramtin Pedarsani, Dimitris Papailiopoulos, and Kannan Ramchandran · 2018
Later among the works it cites.
Communication-computation efficient gradient coding
Min Ye and Emmanuel Abbe · 2018
Later among the works it cites.
A fundamental tradeoff between computation and communication in distributed computing
Songze Li, Mohammad Ali Maddah-Ali, Qian Yu, and A Salman Avestimehr · 2018
Later among the works it cites.
A Unified Coded Deep Neural Network Training Strategy based on Generalized PolyDot codes
Sanghamitra Dutta, Ziqian Bai, Haewon Jeong, Tze Meng Low, and Pulkit Grover · 2018
Later among the works it cites.
Rateless codes for near-perfect load balancing in distributed matrix-vector multiplication
Ankur Mallick, Malhar Chaudhari, and Gauri Joshi · 2018
Later among the works it cites.
An application of storage-optimal matdot codes for coded matrix multiplication: Fast k-nearest neighbors estimation
Utsav Sheth, Sanghamitra Dutta, Malhar Chaudhari, Haewon Jeong, Yaoqing Yang, Jukka Kohonen, Teemu Roos, and Pulkit Grover · 2018
Later among the works it cites.
Adaptive communication strategies to achieve the best error-runtime trade-off in local-update SGD
Jianyu Wang and Gauri Joshi · 2018
Later among the works it cites.
Jianyu Wang and Gauri Joshi · 2018
Later among the works it cites.
Speeding up distributed gradient descent by utilizing non-persistent stragglers
Emre Ozfatura, Deniz Gündüz, and Sennur Ulukus · 2019
Later among the works it cites.
Anytime minibatch with stale gradients
Haider Al-Lawati, Nuwan Ferdinand, and Stark C Draper · 2019
Later among the works it cites.
Robust gradient descent via moment encoding and ldpc codes
Raj Kumar Maity, Ankit Singh Rawa, and Arya Mazumdar · 2019
Later among the works it cites.
Robust and communication-efficient collaborative learning
Amirhossein Reisizadeh, Hossein Taheri, Aryan Mokhtari, Hamed Hassani, and Ramtin Pedarsani · 2019
Later among the works it cites.
Computation scheduling for distributed machine learning with straggling workers
Mohammad Mohammadi Amiri and Deniz Gündüz · 2019
Later among the works it cites.
Federated learning: Challenges, methods, and future directions
Tian Li, Anit Kumar Sahu, Ameet Talwalkar, and Virginia Smith · 2019
Later among the works it cites.
Matcha: Speeding up decentralized sgd via matching decomposition sampling
Jianyu Wang, Anit Kumar Sahu, Zhouyi Yang, Gauri Joshi, and Soummya Kar · 2019
Later among the works it cites.