Fetching the paper…
Reading the bibliography…
Stochastic variance reduced methods have gained a lot of interest recently for empirical risk minimization due to its appealing run time complexity.
Parallel and distributed computation: numerical methods
D. P. Bertsekas and J. N. Tsitsiklis · 1989
Earlier work this paper cites.
Rcv1: A new benchmark collection for text categorization research
D. D. Lewis, Y. Yang, T. G. Rose, and F. Li · 2004
Earlier work this paper cites.
Result analysis of the nips 2003 feature selection challenge
I. Guyon, S. Gunn, A. Ben-Hur, and G. Dror · 2005
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola · 2010
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Y. Zhang, M. J. Wainwright, and J. C. Duchi · 2012
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss minimization
S. Shalev-Shwartz and T. Zhang · 2013
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate Newton-type method
O. Shamir, N. Srebro, and T. Zhang · 2014
Earlier work this paper cites.
A proximal stochastic gradient method with progressive variance reduction
L. Xiao and T. Zhang · 2014
Earlier work this paper cites.
Federated optimization: Distributed optimization beyond the datacenter
J. Konečnỳ, B. McMahan, and D. Ramage · 2015
Earlier work this paper cites.
A universal catalyst for first-order optimization
H. Lin, J. Mairal, and Z. Harchaoui · 2015
Earlier work this paper cites.
DiSCO: Distributed Optimization for Self-Concordant Empirical Loss
Y. Zhang and X. Lin · 2015
Earlier work this paper cites.
Efficient distributed sgd with variance reduction
S. De and T. Goldstein · 2016
Cited alongside, same era.
DSA: Decentralized double stochastic averaging gradient algorithm
A. Mokhtari and A. Ribeiro · 2016
Cited alongside, same era.
Aide: fast and communication efficient distributed optimization
S. J. Reddi, J. Konečnỳ, P. Richtárik, B. Póczós, and A. Smola · 2016
Cited alongside, same era.
Without-replacement sampling for stochastic gradient methods
O. Shamir · 2016
Cited alongside, same era.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic · 2017
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
Z. Allen-Zhu · 2017
Signsgd: Compressed optimisation for non-convex problems
J. Bernstein, Y.-X. Wang, K. Azizzadenesheli, and A. Anandkumar · 2018
Later among the works it cites.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
C. Fang, C. J. Li, Z. Lin, and T. Zhang · 2018
Later among the works it cites.
Dissipativity theory for accelerating stochastic variance reduction: A unified analysis of svrg and katyusha using semidefinite programs
B. Hu, S. Wright, and L. Lessard · 2018
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Y. Lin, S. Han, H. Mao, Y. Wang, and B. Dally · 2018
Later among the works it cites.
The landscape of empirical risk for nonconvex losses
S. Mei, Y. Bai, and A. Montanari · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distributed stochastic variance reduced gradient methods by sampling extra data with replacement
J. D. Lee, Q. Lin, T. Ma, and T. Yang · 2017
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
X. Lian, C. Zhang, H. Zhang, C.-J. Hsieh, W. Zhang, and J. Liu · 2017
Cited alongside, same era.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
L. M. Nguyen, J. Liu, K. Scheinberg, and M. Takáč · 2017
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2017
Cited alongside, same era.
Memory and communication efficient distributed stochastic optimization with minibatch prox
J. Wang, W. Wang, and N. Srebro · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li · 2017
Cited alongside, same era.
V. Smith, S. Forte, M. Chenxin, M. Takáč, M. I. Jordan, and M. Jaggi · 2018
Later among the works it cites.
D 2 {D}^{2} : Decentralized training over decentralized data
H. Tang, X. Lian, M. Yan, C. Zhang, and J. Liu · 2018
Later among the works it cites.
Spiderboost: A class of faster variance-reduced algorithms for nonconvex optimization
Z. Wang, K. Ji, Y. Zhou, Y. Liang, and V. Tarokh · 2018
Later among the works it cites.
Giant: Globally improved approximate newton method for distributed optimization
S. Wang, F. Roosta-Khorasani, P. Xu, and M. W. Mahoney · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
J. Wangni, J. Wang, J. Liu, and T. Zhang · 2018
Later among the works it cites.
A simple stochastic variance reduced algorithm with fast convergence rates
K. Zhou, F. Shang, and J. Cheng · 2018
Later among the works it cites.
Proximal scope for distributed sparse learning
S. Zhao, G.-D. Zhang, M.-W. Li, and W.-J. Li · 2018
Later among the works it cites.
Communication-efficient accurate statistical estimation
J. Fan, Y. Guo, and K. Wang · 2019
Closest in time.
Finite-sum smooth optimization with sarah
L. M. Nguyen, M. van Dijk, D. T. Phan, P. H. Nguyen, T.-W. Weng, and J. R. Kalagnanam · 2019
Closest in time.
Communication-efficient distributed optimization in networks with gradient tracking and variance reduction
B. Li, S. Cen, Y. Chen, and Y. Chi · 2020
Closest in time.