Fetching the paper…
Reading the bibliography…
We consider distributed optimization where the objective function is spread among different devices, each sending incremental model updates to a central server.
Television by pulse code modulation
Goodall, W. M · 1951
Earlier work this paper cites.
A Stochastic Approximation Method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Picture coding using pseudo-random noise
Roberts, L · 1962
Earlier work this paper cites.
Efficient large-scale distributed training of conditional maximum entropy models
McDonald, R., Mohri, M., Silberman, N., Walker, D., and Mann, G. S · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Zinkevich, M., Weimer, M., Li, L., and Smola, A. J · 2010
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chang, C.-C. and Lin, C.-J · 2011
Earlier work this paper cites.
Parallel distributed computing using python
Dalcin, L. D., Paz, R. R., Kler, P. A., and Cosimo, A · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Dekel, O., Gilad-Bachrach, R., Shamir, O., and Xiao, L · 2012
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence _rate for finite training sets
Roux, N. L., Schmidt, M., and Bach, F. R · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss
Shalev-Shwartz, S. and Zhang, T · 2013
Earlier work this paper cites.
Mini-batch primal and dual methods for SVMs
Takáč, M., Bijral, A., Richtárik, P., and Srebro, N · 2013
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S · 2014
Earlier work this paper cites.
Fast distributed coordinate descent for minimizing non-strongly convex losses
Fercoq, O., Qu, Z., Richtárik, P., and Takáč, M · 2014
Earlier work this paper cites.
Communication-efficient distributed dual coordinate ascent
Jaggi, M., Smith, V., Takáč, M., Terhorst, J., Krishnan, S., Hofmann, T., and Jordan, M. I · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate Newton-type method
Shamir, O., Srebro, N., and Zhang, T · 2014
Earlier work this paper cites.
Primal method for ERM with flexible mini-batching schemes and non-convex losses
Csiba, D. and Richtárik, P · 2015
Earlier work this paper cites.
Stochastic dual coordinate ascent with adaptive probabilities
Csiba, D., Qu, Z., and Richtárik, P · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Gupta, S., Agrawal, A., Gopalakrishnan, K., and Narayanan, P · 2015
Earlier work this paper cites.
Variance reduced stochastic gradient descent with neighbors
Hofmann, T., Lucchi, A., Lacoste-Julien, S., and McWilliams, B · 2015
Earlier work this paper cites.
INTERSPEECH 2015, 16th Annual Conference of the International Speech Communication Association, Dresden, Germany, September 6-10, 2015 , 2015. ISCA
Li, H., Meng, H. M., Ma, B., Chng, E., and Xie, L. (eds.) · 2015
Cited alongside, same era.
Adding vs. averaging in distributed primal-dual optimization
Ma, C., Smith, V., Jaggi, M., Jordan, M. I., Richtárik, P., and Takáč, M · 2015
Cited alongside, same era.
Quartz: Randomized dual coordinate ascent with arbitrary sampling
Qu, Z., Richtárik, P., and Zhang, T · 2015
Cited alongside, same era.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2015
Cited alongside, same era.
Scalable distributed DNN training using commodity GPU cloud computing
Strom, N · 2015
Cited alongside, same era.
TernGrad: Ternary gradients to reduce communication in distributed deep learning
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H · 2017
Later among the works it cites.
ImageNet training in 24 minutes
You, Y., Zhang, Z., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2017
Later among the works it cites.
ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning
Zhang, H., Li, J., Kara, K., Alistarh, D., Liu, J., and Zhang, C · 2017
Later among the works it cites.
The convergence of sparsified gradient methods
Alistarh, D., Hoefler, T., Johansson, M., Konstantinov, N., Khirirat, S., and Renggli, C · 2018
Later among the works it cites.
Convex optimization using sparsified stochastic gradient descent with memory
Cordonnier, J.-B · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dryden, N., Moon, T., Jacobs, S. A., and Essen, B. V · 2016
Cited alongside, same era.
Randomized distributed mean estimation: accuracy vs communication
Konečný, J. and Richtárik, P · 2016
Cited alongside, same era.
SDNA: Stochastic dual Newton ascent for empirical risk minimization
Qu, Z., Richtárik, P., Takáč, M., and Fercoq, O · 2016
Cited alongside, same era.
AIDE: Fast and communication efficient distributed optimization
Reddi, S. J., Konečný, J., Richtárik, P., Póczos, B., and Smola, A. J · 2016
Cited alongside, same era.
Distributed coordinate descent method for learning with big data
Richtárik, P. and Takáč, M · 2016
Cited alongside, same era.
CNTK: Microsoft’s open-source deep-learning toolkit
Seide, F. and Agarwal, A · 2016
Cited alongside, same era.
SDCA without duality, regularization, and individual convexity
Shalev-Shwartz, S · 2016
Cited alongside, same era.
Gower, R. M., Richtárik, P., and Bach, F · 2018
Later among the works it cites.
Asynchronous distributed learning with sparse communications and identification
Grishchenko, D., Iutzeler, F., Malick, J., and Amini, M.-R · 2018
Later among the works it cites.
Breaking the span assumption yields fast finite-sum minimization
Hannah, R., Liu, Y., O’Connor, D., and Yin, W · 2018
Later among the works it cites.
Distributed learning with compressed gradients
Khirirat, S., Feyzmahdavian, H. R., and Johansson, M · 2018
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Lin, Y., Han, S., Mao, H., Wang, Y., and Dally, B · 2018
Later among the works it cites.
Distributed learning with compressed gradient differences
Mishchenko, K., Gorbunov, E., Takáč, M., and Richtárik, P · 2018
Later among the works it cites.
SVRG meets SAGA: k-SVRG — a tale of limited memory
Raj, A. and Stich, S. U · 2018
Later among the works it cites.
Local SGD converges fast and communicates little
Stich, S. U · 2018
Later among the works it cites.
Sparsified SGD with memory
Stich, S. U., Cordonnier, J.-B., and Jaggi, M · 2018
Later among the works it cites.
Communication compression for decentralized training
Tang, H., Gan, S., Zhang, C., Zhang, T., and Liu, J · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Wangni, J., Wang, J., Liu, J., and Zhang, T · 2018
Later among the works it cites.
Error compensated quantized SGD and its applications to large-scale distributed optimization
Wu, J., Huang, W., Huang, J., and Zhang, T · 2018
Later among the works it cites.
Direct acceleration of SAGA using sampled negative momentum
Zhou, K · 2018
Later among the works it cites.
Decentralized stochastic optimization and gossip algorithms with compressed communication
Koloskova, A., Stich, S. U., and Jaggi, M · 2019
Closest in time.
Don’t jump through hoops and remove those loops: Svrg and katyusha are better without the outer loop
Kovalev, D., Horváth, S., and Richtárik, P · 2019
Closest in time.
99% of parallel optimization is inevitably a waste of time
Mishchenko, K., Hanzely, F., and Richtárik, P · 2019
Closest in time.