Fetching the paper…
Reading the bibliography…
Stochastic Gradient Descent (SGD) is a popular algorithm that can achieve state-of-the-art performance on a variety of machine learning tasks.
Distributed asynchronous deterministic and stochastic gradient optimization algorithms
J. Tsitsiklis, D. P. Bertsekas, and M. Athans · 1986
Earlier work this paper cites.
Analysis of an approximate gradient projection method with applications to the backpropagation algorithm
Z. Q. Luo and P. Tseng · 1994
Earlier work this paper cites.
Parallel and Distributed Computation: Numerical Methods
D. P. Bertsekas and J. N. Tsitsiklis · 1997
Earlier work this paper cites.
An improved approximation algorithm for multiway cut
G. Călinescu, H. Karloff, and Y. Rabani · 1998
Earlier work this paper cites.
An incremental gradient(-projection) method with momentum term and adaptive stepsize rule
P. Tseng · 1998
Earlier work this paper cites.
Nonlinear Programming
D. P. Bertsekas · 1999
Earlier work this paper cites.
Convergence rate of incremental subgradient algorithms
A. Nedic and D. P. Bertsekas · 2000
Earlier work this paper cites.
An experimental comparison of min-cut/max-flow algorithms for energy minimization in vision
Y. Boykov and V. Kolmogorov · 2004
Earlier work this paper cites.
RCV1: A new benchmark collection for text categorization research
D. Lewis, Y. Yang, T. Rose, and F. Li · 2004
Earlier work this paper cites.
Maximum margin matrix factorization
N. Srebro, J. Rennie, and T. Jaakkola · 2004
Cited alongside, same era.
The landscape of parallel computing research: A view from berkeley
K. Asanovic and · 2006
Cited alongside, same era.
Training linear svms in linear time
T. Joachims · 2006
Cited alongside, same era.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2008
Cited alongside, same era.
MapReduce: simplified data processing on large clusters
J. Dean and S. Ghemawat · 2008
Cited alongside, same era.
SVM Optimization: Inverse dependence on training set size
S. Shalev-Shwartz and N. Srebro · 2008
Cited alongside, same era.
Exact matrix completion via convex optimization
Distributed dual averaging in networks
J. Duchi, A. Agarwal, and M. J. Wainwright · 2010
Later among the works it cites.
Practical large-scale optimization for max-norm regularization
J. Lee, , B. Recht, N. Srebro, R. R. Salakhutdinov, and J. A. Tropp · 2010
Later among the works it cites.
Dremel: Interactive analysis of web-scale datasets
S. Melnik, A. Gubarev, J. J. Long, G. Romer, S. Shivakumar, M. Tolton, and T. Vassilakis · 2010
Later among the works it cites.
Guaranteed minimum rank solutions of matrix equations via nuclear norm minimization
B. Recht, M. Fazel, and P. Parrilo · 2010
Later among the works it cites.
Parallelized stochastic gradient descent
M. Zinkevich, M. Weimer, A. Smola, and L. Li · 2010
Later among the works it cites.
Optimal distributed online prediction using mini-batches
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao · 2011
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Candès and B. Recht · 2009
Cited alongside, same era.
Slow learners are fast
J. Langford, A. J. Smola, and M. Zinkevich · 2009
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Cited alongside, same era.
http://dblife.cs.wisc.edu
A. Doan
Cited in the paper.
https://github.com/JohnLangford/vowpal_wabbit/wiki
J. Langford
Cited in the paper.
The Future of Computing Performance: Game Over or Next Level
S. H. Fuller and L. I. Millett, editors · 2011
Closest in time.
Web scale entity resolution using relational evidence
T. Lee, Z. Wang, H. Wang, and S. Hwang · 2011
Closest in time.
Parallel stochastic gradient algorithms for large-scale matrix completion
B. Recht and C. Ré · 2011
Closest in time.