Fetching the paper…
Reading the bibliography…
State-of-the-art deep learning systems rely on iterative distributed training to tackle the increasing complexity of models and input data.
The Complexity of Flowshop and Jobshop Scheduling
Garey, M. R., Johnson, D. S., and Sethi, R · 1976
Earlier work this paper cites.
Pregel: a system for large-scale graph processing
Malewicz, G., Austern, M. H., Bik, A. J., Dehnert, J. C., Horn, I., Leiser, N., and Czajkowski, G · 2010
Earlier work this paper cites.
Improving the speed of neural networks on CPUs
Vanhoucke, V., Senior, A., and Mao, M. Z · 2011
Earlier work this paper cites.
LFGraph: Simple and fast distributed graph analytics
Hoque, I. and Gupta, I · 2013
Earlier work this paper cites.
Graphx: A resilient distributed graph system on spark
Xin, R. S., Gonzalez, J. E., Franklin, M. J., and Stoica, I · 2013
Earlier work this paper cites.
Exploiting iterative-ness for parallel ML computations
Cui, H., Tumanov, A., Wei, J., Xu, L., Dai, W., Haber-Kucharsky, J., Ho, Q., Ganger, G. R., Gibbons, P. B., Gibson, G. A., et al · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S. E., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2014
Earlier work this paper cites.
Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
Amodei, D., Anubhai, R., Battenberg, E., Case, C., Casper, J., Catanzaro, B., Chen, J., Chrzanowski, M., Coates, A., Diamos, G., Elsen, E., Engel, J., Fan, L., Fougner, C., Han, T., Hannun, A. Y., Jun, B., LeGresley, P., Lin, L., Narang, S., Ng, A. Y., Ozair, S., Prenger, R., Raiman, J., Satheesh, S., Seetapun, D., Sengupta, S., Wang, Y., Wang, Z., Wang, C., Xiao, B., Yogatama, D., Zhan, J., and Zhu, Z · 2015
Earlier work this paper cites.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Chen, T., Li, M., Li, Y., Lin, M., Wang, N., Wang, M., Xiao, T., Xu, B., Zhang, C., and Zhang, Z · 2015
Earlier work this paper cites.
Binaryconnect: Training deep neural networks with binary weights during propagations
Courbariaux, M., Bengio, Y., and David, J.-P · 2015
Cited alongside, same era.
Deep learning with limited numerical precision
Gupta, S., Agrawal, A., Gopalakrishnan, K., and Narayanan, P · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2015
Cited alongside, same era.
Extremely Large Minibatch SGD: Training ResNet-50 on ImageNet in 15 Minutes
Akiba, T., Suzuki, S., and Fukuda, K · 2017
Later among the works it cites.
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Later among the works it cites.
Cho, M., Finkler, U., Kumar, S., Kung, D., Saxena, V., and Sreedhar, D · 2017
Later among the works it cites.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Later among the works it cites.
PyTorch: Tensors and dynamic neural networks in Python with strong GPU acceleration, 2017
Paszke, A., Gross, S., Chintala, S., and Chanan, G · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
TensorFlow: A System for Large-Scale Machine Learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Cited alongside, same era.
An Introduction to Distributed Deep Learning
Arnold, S · 2016
Cited alongside, same era.
GeePS: Scalable deep learning on distributed GPUs with a GPU-specialized parameter server
Cui, H., Zhang, H., Ganger, G. R., Gibbons, P. B., and Xing, E. P · 2016
Cited alongside, same era.
Identity mappings in deep residual networks
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Firecaffe: near-linear acceleration of deep neural network training on compute clusters
Iandola, F. N., Moskewicz, M. W., Ashraf, K., and Keutzer, K · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Cited alongside, same era.
Later among the works it cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H · 2017
Later among the works it cites.
100-epoch ImageNet Training with AlexNet in 24 Minutes
You, Y., Zhang, Z., Hsieh, C., and Demmel, J · 2017
Later among the works it cites.
Poseidon: An Efficient Communication Architecture for Distributed Deep Learning on GPU Clusters
Zhang, H., Zheng, Z., Xu, S., Dai, W., Ho, Q., Liang, X., Hu, Z., Wei, J., Xie, P., and Xing, E. P · 2017
Later among the works it cites.
Distributed tensorflow
Google · 2018
Closest in time.
Horovod: fast and easy distributed deep learning in tensorflow
Sergeev, A. and Balso, M. D · 2018
Closest in time.
On Scale-out Deep Learning Training for Cloud and HPC
Sridharan, S., Vaidyanathan, K., Kalamkar, D., Das, D., Smorkalov, M. E., Shiryaev, M., Mudigere, D., Mellempudi, N., Avancha, S., Kaul, B., et al · 2018
Closest in time.