Fetching the paper…
Reading the bibliography…
Dense Multi-GPU systems have recently gained a lot of attention in the HPC arena.
M. Barnett, L. Shuler, R. van de Geijn, S. Gupta, D. G. Payne, and J. Watts, “Interprocessor collective communication library (InterCom),” in Proceedings of IEEE Scalable High Performance Computing Conference , May 1994, pp. 357–364
1994
Earlier work this paper cites.
M. Shroff and R. A. V. D. Geijn, “CollMark: MPI Collective Communication Benchmark,” Dept. of Computer Sciences, University of Texas at Austin, Tech. Rep., 2000
2000
Earlier work this paper cites.
J. Liu, A. R. Mamidala, and D. K. Panda, “Fast and Scalable MPI-level Broadcast using InfiniBand’s Hardware Multicast Support,” in Parallel and Distributed Processing Symposium, 2004. Proceedings. 18th International , April 2004, p. 10
2004
Earlier work this paper cites.
R. Thakur, R. Rabenseifner, and W. Gropp, “Optimization of Collective Communication Operations in MPICH,” Int. J. High Perform. Comput. Appl. , vol. 19, no. 1, pp. 49–66, Feb. 2005
2005
Earlier work this paper cites.
A. R. Mamidala, L. Chai, H.-W. Jin, and D. K. Panda, “Efficient SMP-aware MPI-level Broadcast over InfiniBand’s Hardware Multicast,” in Proceedings 20th IEEE International Parallel Distributed Processing Symposium , April 2006, p. 8
2006
Earlier work this paper cites.
T. Chiba, T. Endo, and S. Matsuoka, “High-Performance MPI Broadcast Algorithm for Grid Environments Utilizing Multi-lane NICs,” in Seventh IEEE International Symposium on Cluster Computing and the Grid (CCGrid ’07) , May 2007, pp. 487–494
2007
Earlier work this paper cites.
T. Hoefler, C. Siebert, and W. Rehm, “A Practically Constant-time MPI Broadcast Algorithm for Large-scale InfiniBand Clusters with Multicast,” in Proceedings of the 21st IEEE International Parallel & Distributed Processing Symposium (CAC’07 Workshop) , Mar. 2007, p. 232
2007
Earlier work this paper cites.
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. aurelio Ranzato, A. Senior, P. Tucker, K. Yang, Q. V. Le, and A. Y. Ng, “Large scale distributed deep networks,” in Advances in Neural Information Processing Systems 25 , F. Pereira, C. J. C. Burges, L. Bottou, and K. Q. Weinberger, Eds. Curran Associates, Inc., 2012, pp. 1223–1231. [Online]. Available: http://papers.nips.cc/paper/4687-large-scale-distributed-deep-networks.pdf
2012
Earlier work this paper cites.
T. Hoefler, J. Dinan, D. Buntinas, P. Balaji, B. W. Barrett, R. Brightwell, W. Gropp, V. Kale, and R. Thakur, “Leveraging MPI’s One-sided Communication Interface for Shared-memory Programming,” in Proceedings of the 19th European Conference on Recent Advances in the Message Passing Interface , ser. EuroMPI’12. Berlin, Heidelberg: Springer-Verlag, 2012, pp. 132–141
2012
Earlier work this paper cites.
K. Kandalla, A. Venkatesh, K. Hamidouche, S. Potluri, D. Bureddy, and D. K. Panda, “Designing Optimized MPI Broadcast and Allreduce for Many Integrated Core (MIC) InfiniBand Clusters,” in 2013 IEEE 21st Annual Symposium on High-Performance Interconnects , Aug 2013, pp. 63–70
2013
Earlier work this paper cites.
S. Potluri, K. Hamidouche, A. Venkatesh, D. Bureddy, and D. K. Panda, “Efficient Inter-node MPI Communication Using GPUDirect RDMA for InfiniBand Clusters with NVIDIA GPUs,” in Parallel Processing (ICPP), 2013 42nd International Conference on , Oct 2013, pp. 80–89
2013
Earlier work this paper cites.
2014
Cited alongside, same era.
R. Shi, S. Potluri, K. Hamidouche, J. Perkins, M. Li, D. Rossetti, and D. K. Panda, “Designing Efficient Small Message Transfer Mechanism for Inter-node MPI Communication on InfiniBand GPU Clusters,” in 2014 21st International Conference on High Performance Computing (HiPC) , Dec 2014, pp. 1–10
2014
Cited alongside, same era.
2014
Cited alongside, same era.
A. Venkatesh, H. Subramoni, K. Hamidouche, and D. K. Panda, “A High Performance Broadcast Design with Hardware Multicast and GPUDirect RDMA for Streaming Applications on Infiniband Clusters,” in 2014 21st International Conference on High Performance Computing (HiPC) , Dec 2014, pp. 1–10
A. A. Awan, K. Hamidouche, A. Venkatesh, and D. K. Panda, “Efficient Large Message Broadcast using NCCL and CUDA-Aware MPI for Deep Learning,” in Proceedings of the 23rd European MPI Users’ Group Meeting . ACM, 2016, pp. 15–22
2016
Later among the works it cites.
D. S. Banerjee, K. Hamidouche, and D. K. Panda, “Re-Designing CNTK Deep Learning Framework on Modern GPU Enabled Clusters,” in 2016 IEEE International Conference on Cloud Computing Technology and Science (CloudCom) , Dec 2016, pp. 144–151
2016
Later among the works it cites.
2016
Later among the works it cites.
A. A. Awan, K. Hamidouche, J. M. Hashmi, and D. K. Panda, “S-Caffe: Co-designing MPI Runtimes and Caffe for Scalable Deep Learning on Modern GPU Clusters,” in Proceedings of the 22Nd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , ser. PPoPP ’17. ACM, 2017, pp. 193–205
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin et al. , “TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems, 2015,” Software available from tensorflow. org
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, 05 2015. [Online]. Available: http://dx.doi.org/10.1038/nature14539
2015
Cited alongside, same era.
J. Schmidhuber, “Deep Learning in Neural Networks: An Overview,” Neural networks , vol. 61, pp. 85–117, 2015
2015
Cited alongside, same era.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going Deeper with Convolutions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2015, pp. 1–9
2015
Cited alongside, same era.
2015
Cited alongside, same era.
“CNTK,” http://www.cntk.ai/
2016
Cited alongside, same era.
“KESCH: Cray CS-Storm System,” http://www.cscs.ch/computers/kesch_escha/index.html
Cited in the paper.
2017
Closest in time.
C.-H. Chu, X. Lu, A. A. Awan, H. Subramoni, J. Hashmi, B. Elton, and D. K. Panda, “Efficient and Scalable Multi-Source Streaming Broadcast on GPU Clusters for Deep Learning,” in 46th International Conference on Parallel Processing (ICPP-2017) , Aug 2017, [To appear]
2017
Closest in time.
2017
Closest in time.
Cray, “CS-STORM GPU-ACCELERATED CLUSTER SUPERCOMPUTER,” Accessed: August 8, 2026. [Online]. Available: http://www.cray.com/products/computing/cs-series/cs-storm
2026
Closest in time.
NVIDIA, “DGX-1: Essential Instrument of AI Research,” Accessed: August 8, 2026. [Online]. Available: https://www.nvidia.com/en-us/data-center/dgx-1/
2026
Closest in time.
——, “Optimized Primitives for Collective Multi-GPU Communication,” Accessed: August 8, 2026. [Online]. Available: https://github.com/NVIDIA/nccl
2026
Closest in time.
Oak Ridge National Laboratory, “SUMMIT,” Accessed: August 8, 2026. [Online]. Available: https://www.olcf.ornl.gov/summit/
2026
Closest in time.