Fetching the paper…
Reading the bibliography…
We study distributed optimization algorithms for minimizing the average of \emph{heterogeneous} functions distributed across several machines with a focus on communication efficiency.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Mapreduce: simplified data processing on large clusters
Dean, J. and Ghemawat, S · 2008
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L · 2010
Earlier work this paper cites.
Spark: Cluster computing with working sets
Zaharia, M., Chowdhury, M., Franklin, M. J., Shenker, S., and Stoica, I · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Zinkevich, M., Weimer, M., Li, L., and Smola, A. J · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Recht, B., Re, C., Wright, S., and Niu, F · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., et al · 2012
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Dekel, O., Gilad-Bachrach, R., Shamir, O., and Xiao, L · 2012
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Zhang, Y., Wainwright, M. J., and Duchi, J. C · 2012
Earlier work this paper cites.
Estimation, optimization, and parallelism when data is sparse
Duchi, J., Jordan, M. I., and McMahan, B · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Le Roux, N., Schmidt, M. W., and Bach, F. R · 2013
Cited alongside, same era.
Information-theoretic lower bounds for distributed statistical estimation with communication constraints
Zhang, Y., Duchi, J., Jordan, M. I., and Wainwright, M. J · 2013
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S · 2014
Cited alongside, same era.
Communication efficient distributed machine learning with the parameter server
Li, M., Andersen, D. G., Smola, A. J., and Yu, K · 2014
Cited alongside, same era.
Communication-efficient distributed optimization using an approximate newton-type method
Shamir, O., Srebro, N., and Zhang, T · 2014
Cited alongside, same era.
Federated optimization: Distributed machine learning for on-device intelligence
Konečnỳ, J., McMahan, H. B., Ramage, D., and Richtárik, P · 2016
Later among the works it cites.
Aide: Fast and communication efficient distributed optimization
Reddi, S. J., Konečnỳ, J., Richtárik, P., Póczós, B., and Smola, A · 2016
Later among the works it cites.
First-order methods in optimization , volume 25
Beck, A · 2017
Later among the works it cites.
Biased importance sampling for deep neural network training
Katharopoulos, A. and Fleuret, F · 2017
Later among the works it cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A proximal stochastic gradient method with progressive variance reduction
Xiao, L. and Zhang, T · 2014
Cited alongside, same era.
Variance reduction in sgd by distributed importance sampling
Alain, G., Lamb, A., Sankar, C., Courville, A., and Bengio, Y · 2015
Cited alongside, same era.
Concentration inequalities for sampling without replacement
Bardenet, Rémi; Maillard, O.-A · 2015
Cited alongside, same era.
Bouchard, G., Trouillon, T., Perez, J., and Gaidon, A · 2015
Cited alongside, same era.
On variance reduction in stochastic gradient descent and its asynchronous variants
Reddi, S. J., Hefny, A., Sra, S., Poczos, B., and Smola, A. J · 2015
Cited alongside, same era.
Stochastic optimization with importance sampling for regularized loss minimization
Zhao, P. and Zhang, T · 2015
Cited alongside, same era.
Later among the works it cites.
Safe adaptive importance sampling
Stich, S. U., Raj, A., and Jaggi, M · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Later among the works it cites.
Lag: Lazily aggregated gradient for communication-efficient distributed learning
Chen, T., Giannakis, G., Sun, T., and Yin, W · 2018
Later among the works it cites.
Training deep models faster with robust, approximate importance sampling
Johnson, T. B. and Guestrin, C · 2018
Later among the works it cites.
Not all samples are created equal: Deep learning with importance sampling
Katharopoulos, A. and Fleuret, F · 2018
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Later among the works it cites.