Fetching the paper…
Reading the bibliography…
In data-parallel synchronous training of deep neural networks, different devices (replicas) run the same program with different partitions of the training batch, but weight update computation is repeated on all replicas, because the weights do not have a batch dimension to partition.
On the momentum term in gradient descent learning algorithms
Qian, N · 1999
Earlier work this paper cites.
Optimization of Collective Communication Operations in MPICH
Thakur, R., Rabenseifner, R., and Gropp, W · 2005
Earlier work this paper cites.
MPI: A Message-Passing Interface Standard. Version 2.2, September 4th 2009
MPI Forum · 2009
Earlier work this paper cites.
Large Scale Distributed Deep Networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., aurelio Ranzato, M., Senior, A., Tucker, P., Yang, K., Le, Q. V., and Ng, A. Y · 2012
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Li, M., Andersen, D. G., Park, J. W., Smola, A. J., Ahmed, A., Josifovski, V., Long, J., Shekita, E. J., and Su, B.-Y · 2014
Earlier work this paper cites.
Adam: a Method for Stochastic Optimization
Kingma, D. P., and Ba, J. L · 2015
Earlier work this paper cites.
TensorFlow: A System for Large-Scale Machine Learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., Wicke, M., Yu, Y., and Zheng, X · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., Klingner, J., Shah, A., Johnson, M., Liu, X., Łukasz Kaiser, Gouws, S., Kato, Y., Kudo, T., Kazawa, H., Stevens, K., Kurian, G., Patil, N., Wang, W., Young, C., Smith, J., Riesa, J., Rudnick, A., Vinyals, O., Corrado, G., Hughes, M., and Dean, J · 2016
Cited alongside, same era.
Neural collaborative filtering
He, X., Liao, L., Zhang, H., Nie, L., Hu, X., and Chua, T · 2017
Cited alongside, same era.
Attention Is All You Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
https://mlperf.org/training-results-0-6 , 2019
MLPerf Training v0.6 Results · 2019
Later among the works it cites.
https://www.tensorflow.org/xla , 2019
XLA: Optimizing Compiler for TensorFlow · 2019
Later among the works it cites.
BlueConnect: Decomposing All-Reduce for Deep Learning on Heterogeneous Network Hierarchy
Cho, M., Finkler, U., and Kung, D · 2019
Later among the works it cites.
Cloud TPU
Google Cloud · 2019
Later among the works it cites.
Cloud TPU Performance Guide
Google Cloud · 2019
Later among the works it cites.
Using bfloat16 with TensorFlow models
Google Cloud · 2019
Later among the works it cites.
Beyond Data and Model Parallelism for Deep Neural Networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
You, Y., Gitman, I., and Ginsburg, B · 2017
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Chen, D., Lee, H., Ngiam, J., Le, Q. V., and Chen, Z · 2018
Cited alongside, same era.
Mesh-Tensorflow: Deep Learning for Supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., Sepassi, R., and Hechtman, B · 2018
Cited alongside, same era.
Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
Shazeer, N., and Stern, M · 2018
Cited alongside, same era.
Jia, Z., Zaharia, M., and Aiken, A · 2019
Later among the works it cites.
Scale MLPerf-0.6 models on Google TPU-v3 Pods
Kumar, S., Bitorff, V., Chen, D., Chou, C., Hechtman, B., Lee, H., Kumar, N., Mattson, P., Wang, S., Wang, T., and Zhou, Y. X. Z · 2019
Later among the works it cites.