Fetching the paper…
Reading the bibliography…
Consider a number of workers running SGD independently on the same pool of data and averaging the models every once in a while -- a common but not well understood practice.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Rcv1: A new benchmark collection for text categorization research
David D Lewis, Yiming Yang, Tony G Rose, and Fan Li · 2004
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
Tong Zhang · 2004
Earlier work this paper cites.
Parallel stochastic gradient descent
Olivier Delalleau and Yoshua Bengio · 2007
Earlier work this paper cites.
Predicting risk from financial reports with regression
Shimon Kogan, Dimitry Levin, Bryan R Routledge, Jacob S Sagi, and Noah A Smith · 2009
Earlier work this paper cites.
Stochastic convex optimization
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Cited alongside, same era.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Cited alongside, same era.
The million song dataset
Thierry Bertin-Mahieux, Daniel PW Ellis, Brian Whitman, and Paul Lamere · 2011
Cited alongside, same era.
Libsvm dataset page
Chih-Chung Chang and Chih-Jen Lin · 2011
Cited alongside, same era.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Cited alongside, same era.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Cited alongside, same era.
Searching for exotic particles in high-energy physics with deep learning
Pierre Baldi, Peter Sadowski, and Daniel Whiteson · 2014
Later among the works it cites.
Dimmwitted: A study of main-memory statistical analytics
Ce Zhang and Christopher Ré · 2014
Later among the works it cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Later among the works it cites.
Global convergence of stochastic gradient descent for some non-convex matrix problems
Christopher De Sa, Christopher Re, and Kunle Olukotun · 2015
Later among the works it cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Parallel stochastic gradient algorithms for large-scale matrix completion
Benjamin Recht and Christopher Ré · 2013
Cited alongside, same era.
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Later among the works it cites.
Splash: User-friendly programming interface for parallelizing stochastic algorithms
Yuchen Zhang and Michael I Jordan · 2015
Later among the works it cites.
On data dependence in distributed stochastic optimization
Avleen S Bijral, Anand D Sarwate, and Nathan Srebro · 2016
Closest in time.