Fetching the paper…
Reading the bibliography…
Modern learning algorithms use gradient descent updates to train inferential models that best explain data.
G. Ananthanarayanan, S. Kandula, A. G. Greenberg, I. Stoica, Y. Lu, B. Saha, and E. Harris, “Reining in the outliers in map-reduce clusters using mantri,” in OSDI , vol. 10, no. 1, 2010, p. 24
2010
Earlier work this paper cites.
B. Recht, C. Re, S. Wright, and F. Niu, “Hogwild: A lock-free approach to parallelizing stochastic gradient descent,” in Advances in neural information processing systems , 2011, pp. 693–701
2011
Earlier work this paper cites.
R. Gemulla, E. Nijkamp, P. J. Haas, and Y. Sismanis, “Large-scale matrix factorization with distributed stochastic gradient descent,” in Proceedings of the 17th ACM SIGKDD international conference on Knowledge discovery and data mining . ACM, 2011, pp. 69–77
2011
Earlier work this paper cites.
L. D. Dalcin, R. R. Paz, P. A. Kler, and A. Cosimo, “Parallel distributed computing using python,” Advances in Water Resources , vol. 34, no. 9, pp. 1124–1139, 2011
2011
Earlier work this paper cites.
A. Auger and B. Doerr, Theory of randomized search heuristics: Foundations and recent developments . World Scientific, 2011, vol. 1
2011
Earlier work this paper cites.
S. Ross, A First Course in Probability , 9th ed. Pearson, 2012
2012
Earlier work this paper cites.
Y. Zhuang, W.-S. Chin, Y.-C. Juan, and C.-J. Lin, “A fast parallel sgd for matrix factorization in shared memory systems,” in Proceedings of the 7th ACM conference on Recommender systems . ACM, 2013, pp. 249–256
2013
Cited alongside, same era.
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu, “1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns,” in Fifteenth Annual Conference of the International Speech Communication Association , 2014
2014
Cited alongside, same era.
——, “Coded MapReduce,” 53rd Allerton Conference , Sept. 2015
2015
Cited alongside, same era.
R. Tandon, Q. Lei, A. Dimakis, and N. Karampatziakis, “Gradient coding,” NIPS Machine Learning Systems Workshop , 2016
2016
Cited alongside, same era.
S. Dutta, V. Cadambe, and P. Grover, “Short-dot: Computing large linear transforms distributedly using coded short dot products,” in Advances In Neural Information Processing Systems , 2016, pp. 2100–2108
2017
Closest in time.
2017
Closest in time.
K. Lee, M. Lam, R. Pedarsani, D. Papailiopoulos, and K. Ramchandran, “Speeding up distributed machine learning using codes,” IEEE Trans. Inf. Theory , 2017
2017
Closest in time.
S. Li, M. A. Maddah-Ali, Q. Yu, and A. S. Avestimehr, “A fundamental tradeoff between computation and communication in distributed computing,” IEEE Trans. Inf. Theory , 2017
2017
Closest in time.
A. Reisizadeh, S. Prakash, R. Pedarsani, and S. Avestimehr, “Coded computation over heterogeneous clusters,” in IEEE ISIT , 2017, pp. 2408–2412
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
S. Li, M. A. Maddah-Ali, and A. S. Avestimehr, “A unified coding framework for distributed computing with straggling servers,” IEEE NetCod , Sept. 2016
2016
Cited alongside, same era.
2017
Closest in time.