2016

Distributed Deep Learning Using Synchronous Stochastic Gradient Descent

Das, Dipankar, Avancha, Sasikanth, Mudigere, Dheevatsa et al.

Understand

We design and implement a distributed multinode synchronous SGD algorithm, without altering hyper parameters, or compressing data, or altering algorithmic behavior.

  • We perform a detailed analysis of scaling, and identify optimal design points for different networks.
  • We demonstrate scaling of CNNs on 100s of nodes, and present what we believe to be record training throughputs.
  • A 512 minibatch VGG-A CNN training run is scaled 90X on 128 nodes.

Built on

  • ImageNet: A Large-Scale Hierarchical Image Database

    Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009

    Earlier work this paper cites.

  • Conversational speech transcription using context-dependent deep neural networks

    Seide, Frank, Li, Gang, and Yu, Dong · 2011

    Earlier work this paper cites.

  • Overfeat: Integrated recognition, localization and detection using convolutional networks

    Original

    Sermanet, Pierre, Eigen, David, Zhang, Xiang, Mathieu, Michaël, Fergus, Rob, and LeCun, Yann · 2013

    Earlier work this paper cites.

  • Project adam: Building an efficient and scalable deep learning training system

    Chilimbi, Trishul M., Suzue, Yutaka, Apacible, Johnson, and Kalyanaraman, Karthik · 2014

    Earlier work this paper cites.

  • Deep learning in neural networks: An overview

    Original

    Schmidhuber, J · 2014

    Earlier work this paper cites.

Similar

  • 1-bit stochastic gradient descent and application to data-parallel distributed training of speech dnns

    Seide, Frank, Fu, Hao, Droppo, Jasha, Li, Gang, and Yu, Dong · 2014

    Cited alongside, same era.

  • On parallelizability of stochastic gradient descent for speech dnns

    Seide, Frank, Fu, Hao, Droppo, Jasha, Li, Gang, and Yu, Dong · 2014

    Cited alongside, same era.

  • Very deep convolutional networks for large-scale image recognition

    Original

    Simonyan, Karen and Zisserman, Andrew · 2014

    Cited alongside, same era.

  • Deep learning with elastic averaging SGD

    Original

    Zhang, Sixin, Choromanska, Anna, and LeCun, Yann · 2014

    Cited alongside, same era.

Then

  • TensorFlow: Large-scale machine learning on heterogeneous systems, 2015

    Abadi, Martín, Agarwal, Ashish, Barham, Paul, Brevdo, Eugene, Chen, Zhifeng, Citro, Craig, Corrado, Greg S., Davis, Andy, Dean, Jeffrey, Devin, Matthieu, Ghemawat, Sanjay, Goodfellow, Ian, Harp, Andrew, Irving, Geoffrey, Isard, Michael, Jia, Yangqing, Jozefowicz, Rafal, Kaiser, Lukasz, Kudlur, Manjunath, Levenberg, Josh, Mané, Dan, Monga, Rajat, Moore, Sherry, Murray, Derek, Olah, Chris, Schuster, Mike, Shlens, Jonathon, Steiner, Benoit, Sutskever, Ilya, Talwar, Kunal, Tucker, Paul, Vanhoucke, Vincent, Vasudevan, Vijay, Viégas, Fernanda, Vinyals, Oriol, Warden, Pete, Wattenberg, Martin, Wicke, Martin, Yu, Yuan, and Zheng, Xiaoqiang · 2015

    Later among the works it cites.

  • Firecaffe: near-linear acceleration of deep neural network training on compute clusters

    Original

    Iandola, Forrest N., Ashraf, Khalid, Moskewicz, Matthew W., and Keutzer, Kurt · 2015

    Later among the works it cites.

  • Improving concurrency and asynchrony in multithreaded mpi applications using software offloading

    Vaidyanathan, K., Kalamkar, D. D., Pamnany, K., Hammond, J. R., Balaji, P., Das, D., Park, J., and Joo, Balint · 2015

    Later among the works it cites.

  • Deep image: Scaling up image recognition

    Original

    Wu, Ren, Yan, Shengen, Shan, Yi, Dang, Qingqing, and Sun, Gang · 2015

    Later among the works it cites.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…