Fetching the paper…
Reading the bibliography…
This paper proposes a novel parallel stochastic gradient descent (SGD) method that is obtained by applying parallel sets of SGD iterations (each set operating on one node using the data residing in it) for finding the direction in each iteration of a batch descent method.
Cambridge, UK: Cambridge University Press, 2004
S. Boyd and L. Vandenberghe, Convex optimization · 2004
Earlier work this paper cites.
C. Chu, S. Kim, Y. Lin, Y. Yu, G. Bradski, A. Ng, and K. Olukotun, “Map-reduce for machine learning on multicore,” NIPS
2006
Earlier work this paper cites.
C. Hsieh, K. Chang, C. Lin, S. Keerthi, and S. Sundararajan, “A dual coordinate descent method for large-scale linear svm,” in ICML
2008
Earlier work this paper cites.
2008
Earlier work this paper cites.
G. Mann, R. McDonald, M. Mohri, N. Silberman, and D. Walker, “Efficient large-scale distributed training of conditional maximum entropy models,” in NIPS
2009
Cited alongside, same era.
L. Bottou, “Large-scale machine learning with stochastic gradient descent,” in COMPSTAT’2010
2010
Cited alongside, same era.
M. Zinkevich, M. Weimer, A. Smola, and L. Li, “Parallelized stochastic gradient descent,” in NIPS
2010
Cited alongside, same era.
K. Hall, S. Gilpin, and G. Mann, “Mapreduce/bigtable for distributed optimization,” in NIPS Workshop on Leaning on Cores, Clusters, and Clouds
2010
Cited alongside, same era.
A. Agarwal, O. Chapelle, M. Dudik, and J. Langford, “A reliable effective terascale linear learning system,” in arXiv
2011
Later among the works it cites.
N. Le Roux, M. Schmidt, and F. Bach, “A stochastic gradient method with an exponential convergence rate for strongly convex optimization with finite training sets,” in arXiv
2012
Later among the works it cites.
R. Johnson and T. Zhang, “Accelerating stochastic gradient descent using predictive variance reduction,” NIPS
2013
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…