Fetching the paper…
Reading the bibliography…
We introduce ProxSkip -- a surprisingly simple and provably efficient method for minimizing the sum of a smooth ($f$) and an expensive nonsmooth proximable ($\psi$) function.
Backpropagation convergence via deterministic nonmonotone perturbed minimization
Mangasarian, O. L. and Solodov, M. V · 1994
Earlier work this paper cites.
Introductory lectures on convex optimization: a basic course (Applied Optimization)
Nesterov, Y · 2004
Earlier work this paper cites.
Regularization and variable selection via the elastic net
Zhou, H. and Hastie, T · 2005
Earlier work this paper cites.
Proximal splitting methods in signal processing
Combettes, P. L. and Pesquet, J.-C · 2009
Earlier work this paper cites.
Distributed training strategies for the structured perceptron
McDonald, R., Hall, K., and Mann, G · 2010
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chang, C.-C. and Lin, C.-J · 2011
Earlier work this paper cites.
On a generalization of the iterative soft-thresholding algorithm for the case of non-separable penalty
Loris, I. and Verhoeven, C · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Dekel, O., Gilad-Bachrach, R., Shamir, O., and Xiao, L · 2012
Earlier work this paper cites.
An optimal method for stochastic composite optimization
Lan, G · 2012
Earlier work this paper cites.
A primal–dual fixed point algorithm for convex separable minimization with applications to image restoration
Chen, P., Huang, J., and Zhang, X · 2013
Earlier work this paper cites.
Gradient methods for minimizing composite functions
Nesterov, Y · 2013
Earlier work this paper cites.
A forward-backward view of some primal-dual optimization methods in image recovery
Combettes, P. L., Condat, L., Pesquet, J.-C., and Vũ, B. C · 2014
Earlier work this paper cites.
Proximal algorithms
Parikh, N. and Boyd, S · 2014
Earlier work this paper cites.
Understanding machine learning: from theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
Communication complexity of distributed convex learning and optimization
Arjevani, Y. and Shamir, O · 2015
Earlier work this paper cites.
A simple algorithm for a class of nonsmooth convex–concave saddle-point problems
Drori, Y., Sabach, S., and Teboulle, M · 2015
Earlier work this paper cites.
Efficient evaluation of scaled proximal operators
Friedlander, M. and Goh, G · 2016
Earlier work this paper cites.
Federated learning: strategies for improving communication efficiency
Konečný, J., McMahan, H. B., Yu, F., Richtárik, P., Suresh, A. T., and Bacon, D · 2016
Cited alongside, same era.
NEXT: In-network nonconvex optimization
Lorenzo, P. D. and Scutari, G · 2016
Cited alongside, same era.
Federated learning of deep networks using model averaging
McMahan, H. B., Moore, E., Ramage, D., and y Arcas, B. A · 2016
Cited alongside, same era.
Achieving geometric convergence for distributed optimization over time-varying graphs
Nedić, A., Olshevsky, A., and Shi, W · 2016
Cited alongside, same era.
Parallel SGD: When does averaging help?
Zhang, J., De Sa, C., Mitliagkas, I., and Ré, C · 2016
Cited alongside, same era.
Optimal and practical algorithms for smooth and strongly convex decentralized optimization
Kovalev, D., Salim, A., and Richtárik, P · 2020
Later among the works it cites.
Proximal Methods for Image Processing , pp. 165–202
Luke, D. R · 2020
Later among the works it cites.
From local SGD to local fixed-point methods for federated learning
Malinovskiy, G., Kovalev, D., Gasanov, E., Condat, L., and Richtarik, P · 2020
Later among the works it cites.
Dualize, split, randomize: Fast nonsmooth optimization algorithms
Salim, A., Condat, L., Mishchenko, K., and Richtárik, P · 2020
Later among the works it cites.
Federated accelerated stochastic gradient descent
Yuan, H. and Ma, T · 2020
Later among the works it cites.
Generalized monotone operators and their averaged resolvents
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beck, A · 2017
Cited alongside, same era.
Don’t use large mini-batches, use local SGD
Lin, T., Stich, S. U., and Jaggi, M · 2018
Cited alongside, same era.
Ray: A distributed framework for emerging AI applications
Moritz, P., Nishihara, R., Wang, S., Tumanov, A., Liaw, R., Liang, E., Elibol, M., Yang, Z., Paul, W., Jordan, M. I., and Stoica, I · 2018
Cited alongside, same era.
Proximal splitting algorithms for convex optimization: A tour of recent advances, with new twists
Condat, L., Kitahara, D., Contreras, A., and Hirabayashi, A · 2019
Cited alongside, same era.
SGD: General analysis and improved rates
Gower, R. M., Loizou, N., Qian, X., Sailanbayev, A., Shulgin, E., and Richtárik, P · 2019
Cited alongside, same era.
First analysis of local GD on heterogeneous data
Khaled, A., Mishchenko, K., and Richtárik, P · 2019
Cited alongside, same era.
Local SGD converges fast and communicates little
Stich, S. U · 2019
Cited alongside, same era.
Bauschke, H. H., Moursi, W. M., and Wang, X · 2021
Later among the works it cites.
Local SGD: Unified theory and new efficient methods
Gorbunov, E., Hanzely, F., and Richtárik, P · 2021
Later among the works it cites.
Stochastic quasi-gradient methods: Variance reduction via Jacobian sketching
Gower, R. M., Richtárik, P., and Bach, F · 2021
Later among the works it cites.
Advances and open problems in federated learning
Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K., Charles, Z., Cormode, G., Cummings, R., D’Oliveira, R. G. L., Rouayheb, S. E., Evans, D., Gardner, J., Garrett, Z., Gascón, A., Ghazi, B., Gibbons, P. B., Gruteser, M., Harchaoui, Z., He, C., He, L., Huo, Z., Hutchinson, B., Hsu, J., Jaggi, M., Javidi, T., Joshi, G., Khodak, M., Konečný, J., Korolova, A., Koushanfar, F., Koyejo, S., Lepoint, T., Liu, Y., Mittal, P., Mohri, M., Nock, R., Özgür, A., Pagh, R., Raykova, M., Qi, H., Ramage, D., Raskar, R., Song, D., Song, W., Stich, S. U., Sun, Z., Suresh, A. T., Tramèr, F., Vepakomma, P., Wang, J., Xiong, L., Xu, Z., Yang, Q., Yu, F. X., Yu, H., and Zhao, S · 2021
Later among the works it cites.
Breaking the centralized barrier for cross-device federated learning
Karimireddy, S. P., Jaggi, M., Kale, S., Mohri, M., Reddi, S. J., Stich, S. U., and Suresh, A. T · 2021
Later among the works it cites.
An improved analysis of gradient tracking for decentralized machine learning
Koloskova, A., Lin, T., and Stich, S. U · 2021
Later among the works it cites.
Quasi-global momentum: Accelerating decentralized deep learning on heterogeneous data
Lin, T., Karimireddy, S. P., Stich, S., and Jaggi, M · 2021
Later among the works it cites.
Linear convergence in federated learning: Tackling client heterogeneity and sparse gradients
Mitra, A., Jaafar, R., Pappas, G., and Hassani, H · 2021
Later among the works it cites.
Relaysum for decentralized deep learning on heterogeneous data
Vogels, T., He, L., Koloskova, A., Lin, T., Karimireddy, S. P., Stich, S. U., and Jaggi, M · 2021
Later among the works it cites.
The min-max complexity of distributed stochastic convex optimization with intermittent communication
Woodworth, B. E., Bullins, B., Shamir, O., and Srebro, N · 2021
Later among the works it cites.
Removing data heterogeneity influence enhances network topology dependence of decentralized SGD
Yuan, K. and Alghunaim, S. A · 2021
Later among the works it cites.
Distributed proximal splitting algorithms with rates and acceleration
Condat, L., Malinovsky, G., and Richtárik, P · 2022
Closest in time.