Fetching the paper…
Reading the bibliography…
Huge scale machine learning problems are nowadays tackled by distributed optimization algorithms, i.e.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Efficient estimations from a slowly convergent Robbins-Monro process
David Ruppert · 1988
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T. Polyak and Anatoli B. Juditsky · 1992
Earlier work this paper cites.
SciPy: Open source scientific tools for Python, 2001–
Eric Jones, Travis Oliphant, Pearu Peterson, et al · 2001
Earlier work this paper cites.
RCV1: A new benchmark collection for text categorization research
David D. Lewis, Yiming Yang, Tony G. Rose, and Fan Li · 2004
Earlier work this paper cites.
Pascal large scale learning challenge
Soren Sonnenburg, Vojtvech Franc, E. Yom-Tov, and M. Sebag · 2008
Earlier work this paper cites.
On-chip training of recurrent neural networks with limited numerical precision
Taesik Na, Jong Hwan Ko, Jaeha Kung, and Saibal Mukhopadhyay · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R. Bach · 2011
Earlier work this paper cites.
HOGWILD!: A lock-free approach to parallelizing stochastic gradient descent
Feng Niu, Benjamin Recht, Christopher Re, and Stephen J. Wright · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al · 2011
Earlier work this paper cites.
Stochastic Gradient Descent Tricks
Leon Bottou · 2012
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, Martin J Wainwright, and John C Duchi · 2012
Earlier work this paper cites.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Ohad Shamir and Tong Zhang · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Cited alongside, same era.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Cited alongside, same era.
Passcode: Parallel asynchronous stochastic dual co-ordinate descent
Cho-Jui Hsieh, Hsiang-Fu Yu, and Inderjit Dhillon · 2015
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2017
Later among the works it cites.
meProp: Sparsified back propagation for accelerated deep learning with reduced overfitting
Xu Sun, Xuancheng Ren, Shuming Ma, and Houfeng Wang · 2017
Later among the works it cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Later among the works it cites.
Scaling SGD batch size to 32k for ImageNet training
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Later among the works it cites.
ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christopher De Sa, Ce Zhang, Kunle Olukotun, and Christopher Ré · 2015
Cited alongside, same era.
Scalable distributed DNN training using commodity GPU cloud computing
Nikko Strom · 2015
Cited alongside, same era.
Stochastic optimization with importance sampling for regularized loss minimization
Peilin Zhao and Tong Zhang · 2015
Cited alongside, same era.
Communication quantization for data-parallel training of deep neural networks
Nikoli Dryden, Sam Ade Jacobs, Tim Moon, and Brian Van Essen · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Efficient use of limited-memory accelerators for linear learning on heterogeneous systems
Celestine Dünner, Thomas Parnell, and Martin Jaggi · 2017
Cited alongside, same era.
The convergence of stochastic gradient descent in asynchronous shared memory
Dan Alistarh, Christopher De Sa, and Nikola Konstantinov · 2018
Closest in time.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Sarit Khirirat, Nikola Konstantinov, and Cédric Renggli · 2018
Closest in time.
Adacomp : Adaptive residual gradient compression for data-parallel distributed training
Chia-Yu Chen, Jungwook Choi, Daniel Brand, Ankur Agrawal, Wei Zhang, and Kailash Gopalakrishnan · 2018
Closest in time.
Convex optimization using sparsified stochastic gradient descent with memory
Jean-Baptiste Cordonnier · 2018
Closest in time.
Improved asynchronous parallel optimization analysis for stochastic incremental methods
Rémi Leblond, Fabian Pedregosa, and Simon Lacoste-Julien · 2018
Closest in time.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and Bill Dally · 2018
Closest in time.
Local SGD converges fast and communicates little
Sebastian U. Stich · 2018
Closest in time.
Communication compression for decentralized training
Hanlin Tang, Shaoduo Gan, Ce Zhang, and Ji Liu · 2018
Closest in time.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2018
Closest in time.
Error compensated quantized SGD and its applications to large-scale distributed optimization
Jiaxiang Wu, Weidong Huang, Junzhou Huang, and Tong Zhang · 2018
Closest in time.