Fetching the paper…
Reading the bibliography…
Distributed training of massive machine learning models, in particular deep neural networks, via Stochastic Gradient Descent (SGD) is becoming commonplace.
Rcv1: A new benchmark collection for text categorization research
David D Lewis, Yiming Yang, Tony G Rose, and Fan Li · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Project adam: Building an efficient and scalable deep learning training system
Trishul M Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech dnns
F. Seide, H. Fu, L. G. Jasha, and D. Yu · 2014
Earlier work this paper cites.
1-bit Stochastic Gradient Descent and its Application to Data-parallel Distributed Training of Speech DNNs
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu · 2014
Earlier work this paper cites.
An introduction to computational networks and the computational network toolkit
Dong Yu, Adam Eversole, Mike Seltzer, Kaisheng Yao, Zhiheng Huang, Brian Guenter, Oleksii Kuchaiev, Yu Zhang, Frank Seide, Huaming Wang, et al · 2014
Earlier work this paper cites.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Earlier work this paper cites.
Taming the wild: A unified analysis of Hogwild
Christopher De Sa, Ce Zhang, Kunle Olukotun, and Christopher Ré · 2015
Earlier work this paper cites.
Asynchronous stochastic convex optimization
John C Duchi, Sorathan Chaturapruek, and Christopher Ré · 2015
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Earlier work this paper cites.
Asynchronous stochastic coordinate descent: Parallelism and convergence properties
Ji Liu and Stephen J Wright · 2015
Cited alongside, same era.
Privacy-preserving deep learning
Reza Shokri and Vitaly Shmatikov · 2015
Cited alongside, same era.
Scalable distributed dnn training using commodity gpu cloud computing
Nikko Strom · 2015
Cited alongside, same era.
Petuum: A new platform for distributed machine learning on big data
Eric P Xing, Qirong Ho, Wei Dai, Jin Kyu Kim, Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, and Yaoliang Yu · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Communication quantization for data-parallel training of deep neural networks
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J Dally · 2017
Later among the works it cites.
meprop: Sparsified back propagation for accelerated deep learning with reduced overfitting
Xu Sun, Xuancheng Ren, Shuming Ma, and Houfeng Wang · 2017
Later among the works it cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi · 2017
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang · 2017
Later among the works it cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nikoli Dryden, Sam Ade Jacobs, Tim Moon, and Brian Van Essen · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Xnor-net: Imagenet classification using binary convolutional neural networks
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Cited alongside, same era.
QSGD: Randomized quantization for communication-efficient stochastic gradient descent
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Later among the works it cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Later among the works it cites.
Scaling sgd batch size to 32k for imagenet training
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Later among the works it cites.
Yellowfin and the art of momentum tuning
Jian Zhang, Ioannis Mitliagkas, and Christopher Ré · 2017
Later among the works it cites.
The convergence of stochastic gradient descent in asynchronous shared memory
Dan Alistarh, Christopher De Sa, and Nikola Konstantinov · 2018
Closest in time.
Sparcml: High-performance sparse communication for machine learning
Cèdric Renggli, Dan Alistarh, and Torsten Hoefler · 2018
Closest in time.