Fetching the paper…
Reading the bibliography…
Large-scale training is important to ensure high performance and accuracy of machine-learning models.
W. J. Dally, “Performance analysis of k-ary n-cube interconnection networks,” IEEE Transactions on Computers , vol. 39, no. 6, pp. 775–785, 1990
1990
Earlier work this paper cites.
A. Agarwal, “Limits on interconnection network performance,” IEEE Transactions on Parallel and Distributed Systems , vol. 2, no. 4, pp. 398–412, 1991
1991
Earlier work this paper cites.
T. Shanley, Infiniband Network Architecture . Addison-Wesley, 2002
2002
Earlier work this paper cites.
W. Dally and B. Towles, Principles and Practices of Interconnection Networks . San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 2003
2003
Earlier work this paper cites.
M. Al-Fares, A. Loukissas, and A. Vahdat, “A scalable, commodity data center network architecture,” in Proc. 32nd Int. Symp. Computer Architecture (ISCA) , 2005
2005
Earlier work this paper cites.
J. Kim, W. J. Dally, B. Towles, and A. K. Gupta, “Microarchitecture of a high radix router,” in Proc. 32nd Int. Symp. Computer Architecture (ISCA) , 2005
2005
Earlier work this paper cites.
J. Kim, W. J. Dally, S. Scott, and D. Abts, “Technology-driven, highly-scalable dragonfly topology,” in Proc. 35th Annual Int. Symp. Computer Architecture (ISCA) , 2008, pp. 77–88
2008
Earlier work this paper cites.
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola, “Parallelized stochastic gradient descent,” in Advances in Neural Information Processing Systems (NIPS) 23 . Curran Associates, Inc., 2010, pp. 2595–2603
2010
Earlier work this paper cites.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. A. Ranzato, A. Senior, P. Tucker, K. Yang, Q. V. Le, and A. Y. Ng, “Large scale distributed deep networks,” Advances in Neural Information Processing Systems (NIPS) 25 , pp. 1223–1231, 2012
2012
Earlier work this paper cites.
Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun, and O. Temam, “DaDianNao: A machine-learning supercomputer,” in Proc. 47th Annual IEEE/ACM Int. Symp. Microarchitecture (MICRO) , 2014, pp. 609–622
2014
Earlier work this paper cites.
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman, “Project Adam: Building an efficient and scalable deep learning training system,” in Proc. 11th Symp. Operating Systems Design and Implementation (OSDI) 14) , 2014, pp. 571–582
2014
Earlier work this paper cites.
Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell, “Caffe: Convolutional architecture for fast feature embedding,” CoRR , vol. 1408.5093, 2014
2014
Earlier work this paper cites.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015
2015
Earlier work this paper cites.
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang, “MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems,” CoRR , vol. 1512.01274, 2015
2015
Earlier work this paper cites.
L. Bottou, F. E. Curtis, and J. Nocedal, “Optimization methods for large-scale machine learning,” CoRR , vol. 1606.04838, 2016
2016
Earlier work this paper cites.
P. Covington, J. Adams, and E. Sargin, “Deep neural networks for youtube recommendations,” in Proc. 10th ACM Conf. Recommender Systems , 2016, pp. 191–198
2016
Earlier work this paper cites.
D. Das, S. Avancha, D. Mudigere, K. Vaidynathan, S. Sridharan, D. Kalamkar, B. Kaul, and P. Dubey, “Distributed deep learning using synchronous stochastic gradient descent,” CoRR , vol. 1602.06709, 2016
2016
Earlier work this paper cites.
M. Johnson, M. Schuster, Q. V. Le, M. Krikun, Y. Wu, Z. Chen, N. Thorat, F. Viégas, M. Wattenberg, G. Corrado, M. Hughes, and J. Dean, “Google’s multilingual neural machine translation system: enabling zero-shot translation,” CoRR , vol. 1611.04558, 2016
2016
Cited alongside, same era.
D. Foley and J. Danskin, “Ultra-performance Pascal GPU and NVLink interconnect,” Proc. IEEE/ACM Int. Symposium on Microarchitecture (MICRO) , vol. 37, pp. 7–17, 2017
2017
Cited alongside, same era.
A. Gibiansky, “Bringing HPC techniques to deep learning,” Baidu technical blog , 2017
2017
Cited alongside, same era.
B. Brock, Y. Chen, J. Yan, J. D. Owens, A. Buluç, and K. Yelick, “RDMA vs. RPC for implementing distributed data structures,” CoRR , vol. 1910.02158, 2019
2019
Later among the works it cites.
C. Chao and B. Saeta, “Cloud TPU: Codesigning architecture and infrastructure,” in Hot Chips 31 Symposium, Palo Alto, CA, USA , 2019
2019
Later among the works it cites.
U. Gupta, X. Wang, M. Naumov, C. Wu, B. Reagen, D. Brooks, B. Cottel, K. M. Hazelwood, B. Jia, H. S. Lee, A. Malevich, D. Mudigere, M. Smelyanskiy, L. Xiong, and X. Zhang, “The architectural implications of facebook’s dnn-based personalized recommendation,” CoRR , vol. 1906.03109, 2019
2019
Later among the works it cites.
B. Jiang, C. Deng, H. Yi, Z. Hu, G. Zhou, Y. Zheng, S. Huang, X. Guo, D. Wang, Y. Song, and et al., “Xdl: An industrial deep learning framework for high-dimensional sparse data,” in Proc. 1st Int. Workshop on Deep Learning Practice for High-Dimensional Sparse Data , 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P.-l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V. Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D. H. Yoon, “In-datacenter performance analysis of a Tensor Processing Unit,” in Proc. 44th Annual Int. Symp. Computer Architecture (ISCA) , 2017, pp. 1–12
2017
Cited alongside, same era.
X. Pan, J. Chen, R. Monga, S. Bengio, and R. Jozefowicz, “Revisiting distributed synchronous SGD,” 2017
2017
Cited alongside, same era.
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differentiation in PyTorch,” 2017. [Online]. Available: https://pytorch.org/
2017
Cited alongside, same era.
S. Venkataramani, A. Ranjan, S. Banerjee, D. Das, S. Avancha, A. Jagannathan, A. Durg, D. Nagaraj, B. Kaul, P. Dubey, and A. Raghunathan, “ScaleDeep: A scalable compute architecture for learning and evaluating deep networks,” in Proc. 44th Annual Int. Symp. Computer Architecture (ISCA) , 2017, pp. 13–26
2017
Cited alongside, same era.
C. Young, “Evaluation of the Tensor Processing Unit: A deep neural network accelerator for the datacenter,” in Hot Chips 29 Symposium, Palo Alto, CA, USA , 2017
2017
Cited alongside, same era.
“NVidia DGX Pod,” 2018. [Online]. Available: https://www.nvidia.com/en-us/data-center/dgx-pod-reference-architecture/
2018
Cited alongside, same era.
F. Borisyuk, A. Gordo, and V. Sivakumar, “Rosetta: Large scale system for text detection and recognition in images,” in Proc. 24th Int. Conf. Knowledge Discovery & Data Mining (KDD) , 2018, pp. 71–79
2018
Cited alongside, same era.
K. Hazelwood, S. Bird, D. Brooks, S. Chintala, U. Diril, D. Dzhulgakov, M. Fawzy, B. Jia, Y. Jia, A. Kalro, J. Law, K. Lee, J. Lu, P. Noordhuis, M. Smelyanskiy, L. Xiong, and X. Wang, “Applied machine learning at Facebook: A datacenter infrastructure perspective,” in Proc. IEEE Int. Symp. High Performance Computer Architecture (HPCA) , 2018, pp. 620–629
2018
Cited alongside, same era.
R. Mittal, A. Shpiner, A. Panda, E. Zahavi, A. Krishnamurthy, S. Ratnasamy, and S. Shenker, “Revisiting network support for RDMA,” CoRR , vol. 1806.08159, 2018
2018
Cited alongside, same era.
K. Lee, “Introducing Big Basin: Our next-generation AI hardware,” 2019. [Online]. Available: https://engineering.fb.com/data-center-engineering/introducing-big-basin-our-next-generation-ai-hardware
2019
Later among the works it cites.
A. Li, S. L. Song, J. Chen, J. Li, X. Liu, N. Tallent, and K. Barker, “Evaluating modern GPU interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirect,” CoRR , vol. 1903.04611, 2019
2019
Later among the works it cites.
H. Liao, J. Tu, J. Xia, and X. Zhou, “DaVinci: A scalable architecture for neural network computing,” in Hot Chips 31 Symposium, Palo Alto, CA, USA , 2019
2019
Later among the works it cites.
S. Lie, “Wafer scale deep learning,” in Hot Chips 31 Symposium, Palo Alto, CA, USA , 2019
2019
Later among the works it cites.
R. Mayer and H.-A. Jacobsen, “Scalable deep learning on distributed infrastructures: Challenges, techniques and tools,” ACM Computing Surveys , vol. 53, 2019
2019
Later among the works it cites.
E. Medina, “Habana labs approach to scaling AI training,” in Hot Chips 31 Symposium, Palo Alto, CA, USA , 2019
2019
Later among the works it cites.
M. Naumov, D. Mudigere, H. M. Shi, J. Huang, N. Sundaraman, J. Park, X. Wang, U. Gupta, C. Wu, A. G. Azzolini, D. Dzhulgakov, A. Mallevich, I. Cherniavskii, Y. Lu, R. Krishnamoorthi, A. Yu, V. Kondratenko, S. Pereira, X. Chen, W. Chen, V. Rao, B. Jia, L. Xiong, and M. Smelyanskiy, “Deep learning recommendation model for personalization and recommendation systems,” CoRR , vol. 1906.00091, 2019. [Online]. Available: https://github.com/facebookresearch/dlrm
2019
Later among the works it cites.
M. Smelyanskiy, “Zion: Facebook next-generation large-memory unified training platform,” in Hot Chips 31 Symposium, Palo Alto, CA, USA , 2019
2019
Later among the works it cites.
M. Wang, C. Meng, G. Long, C. Wu, J. Yang, W. Lin, and Y. Jia, “Characterizing deep learning training workloads on Alibaba-PAI,” CoRR , vol. 1910.05930, 2019
2019
Later among the works it cites.
A. Yang, N. Garegrat, C. Miao, and K. Vaidyanathan, “Deep learning training at scale: Spring Crest deep learning accelerator,” in Hot Chips 31 Symposium, Palo Alto, CA, USA , 2019
2019
Later among the works it cites.
W. Zhao, “OCP Accelerator Module (OAM),” 2019. [Online]. Available: https://engineering.fb.com/data-center-engineering/accelerator-modules/
2019
Later among the works it cites.
J. Dong, Z. Cao, T. Zhang, J. Ye, S. Wang, F. Feng, L. Zhao, X. Liu, L. Song, L. Peng, Y. Guo, X. Jiang, L. Tang, Y. Du, Y. Zhang, P. Pan, and Y. Xie, “EFLOPS: Algorithm and system co-design for a high performance distributed training platform,” Proc. IEEE Int. Symp. High-Performance Computer Architecture (HPCA) , 2020
2020
Closest in time.
Q. Zheng, B.-Y. Su, J. Yang, A. Azzolini, Q. Wu, O. Jin, S. Karandikar, H. Lupesko, L. Xiong, and E. Zhou, “Shadowsync: Performing synchronization in the background for highly scalable distributed training,” CoRR , vol. 2003.03477, 2020
2020
Closest in time.