Fetching the paper…
Reading the bibliography…
We propose a generic variance-reduced algorithm, which we call MUltiple RANdomized Algorithm (MURANA), for minimizing a sum of several smooth functions plus a regularizer, in a sequential or distributed manner.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Earlier work this paper cites.
Linear convergence with condition number independent access of full gradients
L. Zhang, M. Mahdavi, and R. Jin · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Earlier work this paper cites.
Proximal algorithms
N. Parikh and S. Boyd · 2014
Earlier work this paper cites.
A proximal stochastic gradient method with progressive variance reduction
L. Xiao and T. Zhang · 2014
Earlier work this paper cites.
Variance reduced stochastic gradient descent with neighbors
T. Hofmann, A. Lucchi, S. Lacoste-Julien, and B. McWilliams · 2015
Earlier work this paper cites.
Federated learning: Strategies for improving communication efficiency
J. Konečný, H. B. McMahan, F. X. Yu, P. Richtárik, A. T. Suresh, and D. Bacon · 2016
Earlier work this paper cites.
Parallel coordinate descent methods for big data optimization
P. Richtárik and M. Takáč · 2016
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic · 2017
Earlier work this paper cites.
Convex Analysis and Monotone Operator Theory in Hilbert Spaces
H. H. Bauschke and P. L. Combettes · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
H. Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas · 2017
Earlier work this paper cites.
TernGrad: Ternary gradients to reduce communication in distributed deep learning
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li · 2017
Earlier work this paper cites.
Gradient sparsification for communication-efficient distributed optimization
J. Wangni, J. Wang, J. Liu, and T. Zhang · 2018
Earlier work this paper cites.
Optimal mini-batch and step sizes for SAGA
N. Gazagnadou, R. Gower, and J. Salmon · 2019
Cited alongside, same era.
SGD: General analysis and improved rates
R. M. Gower, N. Loizou, X. Qian, A. Sailanbayev, E. Shulgin, and P. Richtárik · 2019
Cited alongside, same era.
Stochastic distributed learning with gradient quantization and variance reduction
S. Horváth, D. Kovalev, K. Mishchenko, S. Stich, and P. Richtárik · 2019
Cited alongside, same era.
Distributed learning with compressed gradient differences
K. Mishchenko, E. Gorbunov, M. Takáč, and P. Richtárik · 2019
Cited alongside, same era.
MISO is making a comeback with better proofs and rates
X. Qian, A. Sailanbayev, K. Mishchenko, and P. Richtárik · 2019
Cited alongside, same era.
Don’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop
D. Kovalev, S. Horváth, and P. Richtárik · 2020
Later among the works it cites.
Federated learning: Challenges, methods, and future directions
T. Li, A. K. Sahu, A. Talwalkar, and V. Smith · 2020
Later among the works it cites.
A double residual compression algorithm for efficient distributed learning
X. Liu, Y. Li, J. Tang, and M. Yan · 2020
Later among the works it cites.
C. Philippenko and A. Dieuleveut · 2020
Later among the works it cites.
Robust and communication-efficient federated learning from non-i.i.d. data
F. Sattler, S. Wiedemann, K.-R. Müller, and W. Samek · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards closing the gap between the theory and practice of SVRG
O. Sebbouh, N. Gazagnadou, S. Jelassi, F. Bach, and R. Gower · 2019
Cited alongside, same era.
Doublesqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression
H. Tang, C. Yu, X. Lian, T. Zhang, and J. Liu · 2019
Cited alongside, same era.
Optimal gradient compression for distributed and federated learning
A. Albasyoni, M. Safaryan, L. Condat, and P. Richtárik · 2020
Cited alongside, same era.
Qsparse-Local-SGD: Distributed SGD With Quantization, Sparsification, and Local Computations
D. Basu, D. Data, C. Karakus, and S. N. Diggavi · 2020
Cited alongside, same era.
On the discrepancy between the theoretical analysis and practical implementations of compressed communication for distributed deep learning
A. Dutta, E. H. Bergou, A. M. Abdelmoniem, C. Y. Ho, A. N. Sahu, M. Canini, and P. Kalnis · 2020
Cited alongside, same era.
Variance-reduced methods for machine learning
R. M. Gower, M. Schmidt, F. Bach, and P. Richtárik · 2020
Cited alongside, same era.
Unified analysis of stochastic gradient methods for composite convex and smooth optimization
A. Khaled, O. Sebbouh, N. Loizou, R. M. Gower M., and P. Richtárik · 2020
Cited alongside, same era.
Learning theory from first principles
F. Bach · 2021
Closest in time.
Fixed point strategies in data science
P. L. Combettes and J.-C. Pesquet · 2021
Closest in time.
Stochastic quasi-gradient methods: Variance reduction via Jacobian sketching
R. M. Gower, P. Richtárik, and F. Bach · 2021
Closest in time.
Advances and open problems in federated learning
P. Kairouz et al · 2021
Closest in time.
L-SVRG and L-Katyusha with arbitrary sampling
X. Qian, Z. Qu, and P. Richtárik · 2021
Closest in time.
GRACE: A compressed communication framework for distributed machine learning
H. Xu, C.-Y. Ho, A. M. Abdelmoniem, A. Dutta, E. H. Bergou, K. Karatsenidis, M. Canini, and P. Kalnis · 2021
Closest in time.
Dualize, split, randomize: Fast nonsmooth optimization algorithms
A. Salim, L. Condat, K. Mishchenko, and P. Richtárik · 2022
Closest in time.