The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and Cédric Renggli · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
Original
L. Bottou, F. Curtis, and J. Nocedal · 2018
Later among the works it cites.
Stochastic gradient descent with biased but consistent gradient estimators
Original
Jie Chen and Ronny Luss · 2018
Later among the works it cites.
Adaptive balancing of gradient and update computation times using global geometry and approximate subproblems
Sai Praneeth Reddy Karimireddy, Sebastian Stich, and Martin Jaggi · 2018
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Later among the works it cites.
Zeroth-order stochastic variance reduction for nonconvex optimization
Sijia Liu, Bhavya Kailkhura, Pin-Yu Chen, Paishun Ting, Shiyu Chang, and Lisa Amini · 2018
Later among the works it cites.
Sparsified SGD with memory
Sebastian U. Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Later among the works it cites.
ATOMO: communication-efficient learning via atomic sparsification
Hongyi Wang, Scott Sievert, Shengchao Liu, Zachary B. Charles, Dimitris S. Papailiopoulos, and Stephen Wright · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Later among the works it cites.
Efficient greedy coordinate descent for composite problems
Sai Praneeth Karimireddy, Anastasia Koloskova, Sebastian U. Stich, and Martin Jaggi · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Unified optimal analysis of the (stochastic) gradient method
Original
Sebastian U. Stich · 2019
Later among the works it cites.
The error-feedback framework: Better rates for SGD with delayed gradients and compressed communication
Original
Sebastian U. Stich and Sai Praneeth Karimireddy · 2019
Later among the works it cites.
On biased compression for distributed learning
Original
Aleksandr Beznosikov, Samuel Horváth, Peter Richtárik, and Mher Safaryan · 2020
Closest in time.
Finite-time error bounds for biased stochastic approximation with applications to q-learning
Gang Wang and Georgios B. Giannakis · 2020
Closest in time.