Fetching the paper…
Reading the bibliography…
Distributed optimization algorithms are widely used in many industrial machine learning applications.
A bridging model for parallel computation
L. G. Valiant · 1990
Earlier work this paper cites.
Vowpal wabbit online learning project, 2007
J. Langford, L. Li, and A. Strehl · 2007
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Systemml: Declarative machine learning on mapreduce
A. Ghoting, R. Krishnamurthy, E. Pednault, B. Reinwald, V. Sindhwani, S. Tatikonda, Y. Tian, and S. Vaithyanathan · 2011
Earlier work this paper cites.
No one (cluster) size fits all: automatic cluster sizing for data-intensive analytics
H. Herodotou, F. Dong, and S. Babu · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. A. Zinkevich, M. Weimer, et al · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao · 2012
Earlier work this paper cites.
Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing
M. Zaharia, M. Chowdhury, T. Das, A. Dave, J. Ma, M. McCauley, M. J. Franklin, S. Shenker, and I. Stoica · 2012
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss minimization
S. Shalev-Shwartz and T. Zhang · 2013
Earlier work this paper cites.
Trading computation for communication: Distributed stochastic dual coordinate ascent
T. Yang · 2013
Cited alongside, same era.
Communication-efficient distributed dual coordinate ascent
M. Jaggi, V. Smith, et al · 2014
Cited alongside, same era.
Communication efficient distributed machine learning with the parameter server
M. Li, D. G. Andersen, et al · 2014
Cited alongside, same era.
Efficient mini-batch training for stochastic optimization
M. Li, T. Zhang, Y. Chen, and A. J. Smola · 2014
Cited alongside, same era.
Res: Regularized stochastic bfgs algorithm
A. Mokhtari and A. Ribeiro · 2014
Cited alongside, same era.
Dimmwitted: A study of main-memory statistical analytics
C. Zhang and C. Ré · 2014
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
M. Andrychowicz, M. Denil, S. Gomez, M. W. Hoffman, D. Pfau, T. Schaul, and N. de Freitas · 2016
Later among the works it cites.
Designing neural network architectures using reinforcement learning
B. Baker, O. Gupta, N. Naik, and R. Raskar · 2016
Later among the works it cites.
Learning to learn for global optimization of black box functions
Y. Chen, M. W. Hoffman, S. G. Colmenarejo, M. Denil, T. P. Lillicrap, and N. de Freitas · 2016
Later among the works it cites.
RL ^ 2 \hat{~}2 : Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Later among the works it cites.
Omnivore: An optimizer for multi-device deep learning on cpus and gpus
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, et al · 2015
Cited alongside, same era.
MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, , and Z. Zhang · 2015
Cited alongside, same era.
Adding vs. averaging in distributed primal-dual optimization
C. Ma, V. Smith, et al · 2015
Cited alongside, same era.
Sdna: stochastic dual newton ascent for empirical risk minimization
Z. Qu, P. Richtárik, M. Takáč, and O. Fercoq · 2015
Cited alongside, same era.
Splash: User-friendly Programming Interface for Parallelizing Stochastic Algorithms
Y. Zhang and M. I. Jordan · 2015
Cited alongside, same era.
Learning to compose neural networks for question answering
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Cited alongside, same era.
S. Hadjis, C. Zhang, I. Mitliagkas, D. Iter, and C. Re · 2016
Later among the works it cites.
STRADS: A Distributed Framework for Scheduled Model Parallel Machine Learning
J. K. Kim, Q. Ho, S. Lee, X. Zheng, W. Dai, G. A. Gibson, and E. P. Xing · 2016
Later among the works it cites.
Mllib: Machine learning in apache spark
X. Meng, J. Bradley, B. Yuvaz, E. Sparks, S. Venkataraman, D. Liu, J. Freeman, D. Tsai, M. Amde, S. Owen, et al · 2016
Later among the works it cites.
Asynchrony begets momentum, with an application to deep learning
I. Mitliagkas, C. Zhang, S. Hadjis, and C. Re · 2016
Later among the works it cites.
A linearly-convergent stochastic l-bfgs algorithm
P. Moritz, R. Nishihara, and M. I. Jordan · 2016
Later among the works it cites.
Keystoneml: Optimizing pipelines for large-scale advanced analytics
E. R. Sparks, S. Venkataraman, T. Kaftan, M. J. Franklin, and B. Recht · 2016
Later among the works it cites.
Ernest: Efficient Performance Prediction for Large-Scale Advanced Analytics
S. Venkataraman, Z. Yang, M. Franklin, B. Recht, and I. Stoica · 2016
Later among the works it cites.