Fetching the paper…
Reading the bibliography…
Due to its efficiency and ease to implement, stochastic gradient descent (SGD) has been widely used in machine learning.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Francis R. Bach and Eric Moulines · 2011
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
Stochastic dual coordinate ascent methods for regularized loss
Shai Shalev-Shwartz and Tong Zhang · 2013
Earlier work this paper cites.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Ohad Shamir and Tong Zhang · 2013
Earlier work this paper cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis R. Bach, and Simon Lacoste-Julien · 2014
Earlier work this paper cites.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
Elad Hazan and Satyen Kale · 2014
Earlier work this paper cites.
Efficient mini-batch training for stochastic optimization
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J. Smola · 2014
Earlier work this paper cites.
Compressing neural networks with the hashing trick
Wenlin Chen, James T. Wilson, Stephen Tyree, Kilian Q. Weinberger, and Yixin Chen · 2015
Cited alongside, same era.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Cited alongside, same era.
Xnor-net: Imagenet classification using binary convolutional neural networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Minimizing finite sums with the stochastic average gradient
Mark W. Schmidt, Nicolas Le Roux, and Francis R. Bach · 2017
Later among the works it cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Later among the works it cites.
Stochastic convex optimization: Faster local growth implies faster global convergence
Yi Xu, Qihang Lin, and Tianbao Yang · 2017
Later among the works it cites.
Zipml: Training linear models with end-to-end low precision, and a little bit of deep learning
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alham Fikri Aji and Kenneth Heafield · 2017
Cited alongside, same era.
QSGD: communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Christopher De Sa, Megan Leszczynski, Jian Zhang, Alana Marzoev, Christopher R. Aberger, Kunle Olukotun, and Christopher Ré · 2018
Later among the works it cites.
Decentralization meets quantization
Hanlin Tang, Ce Zhang, Shaoduo Gan, Tong Zhang, and Ji Liu · 2018
Later among the works it cites.