Fetching the paper…
Reading the bibliography…
Data parallel training is widely used for scaling distributed deep neural network (DNN) training.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Linux Traffic Control, 1999
Alexey N. Kuznetsov · 1999
Earlier work this paper cites.
Gigabit Ethernet Networking
Cunningham, D., Lane, B., and Lane, W · 1999
Earlier work this paper cites.
Infiniband
Shanley, T · 2002
Earlier work this paper cites.
Fpc: A high-speed compressor for double-precision floating-point data
Burtscher, M. and Ratanaworabhan, P · 2008
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L · 2010
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., and Darrell, T · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Li, M., Andersen, D. G., Park, J. W., Smola, A. J., Ahmed, A., Josifovski, V., Long, J., Shekita, E. J., and Su, B.-Y · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech dnns
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Chen, T., Li, M., Li, Y., Lin, M., Wang, N., Wang, M., Xiao, T., Xu, B., Zhang, C., and Zhang, Z · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., Wicke, M., Yu, Y., and Zheng, X · 2016
Cited alongside, same era.
Revisiting distributed synchronous SGD
Chen, J., Monga, R., Bengio, S., and Józefowicz, R · 2016
Cited alongside, same era.
Distributed training of deep neural networks: Theoretical and practical limits of parallel scalability
Keuper, J. and Preundt, F.-J · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Aji, A. F. and Heafield, K · 2017
Cited alongside, same era.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Network evolution for dnns
Alan, M., Panda, A., Bottini, D., Jian, L., Kumar, P., and Shenker, S · 2018
Later among the works it cites.
Gossipgrad: Scalable deep learning using gossip communication based asynchronous gradient descent
Daily, J., Vishnu, A., Siegel, C., Warfel, T., and Amatya, V · 2018
Later among the works it cites.
DNN model compression under accuracy constraints, 2018
Khoram, S. and Li, J · 2018
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Lin, Y., Han, S., Mao, H., Wang, Y., and Dally, B · 2018
Later among the works it cites.
Parameter hub: a rack-scale parameter server for distributed deep neural network training
Luo, L., Nelson, J., Ceze, L., Phanishayee, A., and Krishnamurthy, A · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Optimized broadcast for deep learning workloads on dense-gpu infiniband clusters: MPI or nccl?
Awan, A. A., Chu, C., Subramoni, H., and Panda, D. K · 2017
Cited alongside, same era.
Adacomp : Adaptive residual gradient compression for data-parallel distributed training
Chen, C., Choi, J., Brand, D., Agrawal, A., Zhang, W., and Gopalakrishnan, K · 2017
Cited alongside, same era.
Sockeye: A toolkit for neural machine translation
Hieber, F., Domhan, T., Denkowski, M., Vilar, D., Sokolov, A., Clifton, A., and Post, M · 2017
Cited alongside, same era.
Nvidia Quadro P4000, 2017
Nvidia · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Cited alongside, same era.
Performance modeling and evaluation of distributed deep learning frameworks on gpus
Shi, S. and Chu, X · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H · 2017
Cited alongside, same era.
InfiniBand Cards, 2018
Mellanox · 2018
Later among the works it cites.
DNN training benchmarks, 2018
MLPerf · 2018
Later among the works it cites.
cuda c programming guide, 2018
Nvidia · 2018
Later among the works it cites.
TBD: benchmarking and analyzing deep neural network training
Zhu, H., Akrout, M., Zheng, B., Pelegris, A., Phanishayee, A., Schroeder, B., and Pekhimenko, G · 2018
Later among the works it cites.
Amazon Web Services, 2019
Amazon · 2019
Closest in time.
Easy Amazon EC2 instance comparison, 2019
ec2instances.info · 2019
Closest in time.
Bandwidth Monitor NG (Next Generation), 2019
Volker Gropp · 2019
Closest in time.