Fetching the paper…
Reading the bibliography…
Distributed data-parallel algorithms aim to accelerate the training of deep neural networks by parallelizing the computation of large mini-batch gradient updates across multiple nodes.
Products of indecomposible, aperiodic, stochastic matrices
Wolfowitz, J · 1963
Earlier work this paper cites.
Non-negative Matrices and Markov Chains
Seneta, E · 1981
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Gossip-based computation of aggregate information
Kempe, D., Dobra, A., and Gehrke, J · 2003
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J. Z., Tomioka, R., and Vojnovic, M · 2007
Earlier work this paper cites.
TernGrad: Ternary gradients to reduce communication in distributed deep learning
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H · 2007
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Senior, A., Tucker, P., Yang, K., Le, Q. V., et al · 2012
Earlier work this paper cites.
Average consensus in the presence of delays in directed graph topologies
Hadjicostis, C. N. and Charalambous, T · 2014
Earlier work this paper cites.
Fast distributed gradient methods
Jakovetić, D., Xavier, J., and Moura, J. M · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Li, M., Andersen, D. G., Park, J. W., Smola, A. J., Ahmed, A., Josifovski, V., Long, J., Shekita, E. J., and Su, B.-Y · 2014
Earlier work this paper cites.
Distributed finite-time average consensus in digraphs in the presence of time delays
Charalambous, T., Yuan, Y., Yang, T., Pan, W., Hadjicostis, C. N., and Johansson, M · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Cited alongside, same era.
Deep learning with elastic averaged SGD
Zhang, S., Choromanska, A., and LeCun, Y · 2015
Cited alongside, same era.
Gossip training for deep learning
Blot, M., Picard, D., Cord, M., and Thome, N · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
How to scale distributed deep learning?
Jin, P. H., Yuan, Q., Iandola, F., and Keutzer, K · 2016
Cited alongside, same era.
Stochastic gradient-push for strongly convex functions on time-varying directed graphs
Nedić, A. and Olshevsky, A · 2016
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Lian, X., Zhang, C., Zhang, H., Hsieh, C.-J., Zhang, W., and Liu, J · 2017
Later among the works it cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, H. B., Moore, E., Ramage, D., Hampson, S., and Agüera y Arcas, B · 2017
Later among the works it cites.
Pytorch: Tensors and dynamic neural networks in python with strong gpu acceleration, 2017
Paszke, A., Chintala, S., Collobert, R., Kavukcuoglu, K., Farabet, C., Bengio, S., Melvin, I., Weston, J., and Mariethoz, J · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Later among the works it cites.
Assran, M. and Rabbat, M · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et al · 2016
Cited alongside, same era.
Extremely large minibatch sgd: training resnet-50 on imagenet in 15 minutes
Akiba, T., Suzuki, S., and Fukuda, K · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
Gehring, J., Auli, M., Grangier, D., Yarats, D., and Dauphin, Y. N · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Cited alongside, same era.
Collaborative deep learning in fixed topology networks
Jiang, Z., Balu, A., Hegde, C., and Sarkar, S · 2017
Cited alongside, same era.
Jia, X., Song, S., He, W., Wang, Y., Rong, H., Zhou, F., Xie, L., Guo, Z., Yang, Y., Yu, L., et al · 2018
Closest in time.
Asynchronous decentralized parallel stochastic gradient descent
Lian, X., Zhang, W., Zhang, C., and Liu, J · 2018
Closest in time.
Exploring the limits of weakly supervised pretraining
Mahajan, D., Girshick, R., Ramanathan, V., He, K., Paluri, M., Li, Y., Bharambe, A., and van der Maaten, L · 2018
Closest in time.
Network topology and communication-computation tradeoffs in decentralized optimization
Nedić, A., Olshevsky, A., and Rabbat, M. G · 2018
Closest in time.
Scaling neural machine translation
Ott, M., Edunov, Sergey, G. D., and Auli, M · 2018
Closest in time.
signSGD with majority vote is communication efficient and fault tolerant
Bernstein, J., Zhao, J., Azizzadenesheli, K., and Anandkumar, A · 2019
Closest in time.