Fetching the paper…
Reading the bibliography…
It is safe to assume that, for the foreseeable future, machine learning, especially deep learning will remain both data- and computation-hungry.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Cited alongside, same era.
REM: Resource-efficient mining for blockchains
Fan Zhang, Ittay Eyal, Robert Escriva, Ari Juels, and Robbert Van Renesse
Cited in the paper.
ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang
Cited in the paper.
Qsgd: Randomized quantization for communication-optimal stochastic gradient descent
Dan Alistarh, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…