Distributed gradient descent (DGD) is an efficient way of implementing gradient descent (GD), especially for large data sets, by dividing the computation tasks into smaller subtasks and assigning to different computing servers (CSs) to be executed in parallel.
In standard parallel execution, per-iteration waiting time is limited by the execution time of the straggling servers.
Coded DGD techniques have been introduced recently, which can tolerate straggling servers via assigning redundant computation tasks to the CSs.
In most of the existing DGD schemes, either with coded computation or coded communication, the non-straggling CSs transmit one message per iteration once they complete all their assigned computation tasks.
Built on
R. Tandon, Q. Lei, A. G. Dimakis, and N. Karampatziakis, “Gradient coding: Avoiding stragglers in distributed learning,” in Proceedings of the 34th International Conference on Machine Learning , ser. Proc. Machine Learning Research, D. Precup and Y. W. Teh, Eds., vol. 70, Sydney, Australia, Aug. 2017, pp. 3368–3376
N. Ferdinand, B. Gharachorloo, and S. C. Draper, “Anytime exploitation of stragglers in synchronous stochastic gradient descent,” in IEEE Int’l Conf. on Machine Learning and Applications (ICMLA) , Dec. 2017, pp. 141–146
2017
Earlier work this paper cites.
K. Lee, M. Lam, R. Pedarsani, D. Papailiopoulos, and K. Ramchandran, “Speeding up distributed machine learning using codes,” IEEE Trans. on Information Theory , vol. 64, no. 3, pp. 1514–1529, Mar. 2018
2018
Earlier work this paper cites.
Similar
S. Dutta, G. Joshi, S. Ghosh, P. Dube, and P. Nagpurkar, “Slow and stale gradients can win the race: Error-runtime trade-offs in distributed SGD,” in The 21st International Conference on Artificial Intelligence and Statistics (AISTATS) , 2018
2018
Cited alongside, same era.
N. Ferdinand and S. C. Draper, “Hierarchical coded computation,” in IEEE Int’l Symp. on Information Theory (ISIT) , Vail, CO, Jun. 2018
2018
Cited alongside, same era.
R. K. Maity, A. S. Rawat, and A. Mazumdar, “Robust gradient descent via moment encoding with ldpc codes,” SysML Conference , 2018
A. Mallick, M. Chaudhari, and G. Joshi, “Rateless codes for near-perfect load balancing in distributed matrix-vector multiplication,” CoRR , vol. abs/1804.10331, 2018. [Online]. Available: http://arxiv.org/abs/1804.10331