Fetching the paper…
Reading the bibliography…
High performance multi-GPU computing becomes an inevitable trend due to the ever-increasing demand on computation capability in emerging domains such as deep learning, big data and planet-scale simulations.
J. D. Little, “A proof for the queuing formula: L= λ \lambda W,” Operations research , 1961
1961
Earlier work this paper cites.
NVIDIA, “SLI best practices,” Tech. Rep., 2007
2007
Earlier work this paper cites.
AMD, “ATI CrossFire Pro User Guide,” Tech. Rep., 2009
2009
Earlier work this paper cites.
S. Pabst, A. Koch, and W. Straßer, “Fast and scalable cpu/gpu collision detection for rigid and deformable surfaces,” in Computer Graphics Forum . Wiley Online Library, 2010
2010
Earlier work this paper cites.
D. Ziakas, A. Baum, R. A. Maddox, and R. J. Safranek, “Intel® quickpath interconnect architectural features supporting scalable system architectures,” in High Performance Interconnects (HOTI), 2010 IEEE 18th Annual Symposium on . IEEE, 2010
2010
Earlier work this paper cites.
P. Grun, “Introduction to infiniband for end users,” White paper, InfiniBand Trade Association , 2010
2010
Earlier work this paper cites.
K. Spafford, J. S. Meredith, and J. S. Vetter, “Quantifying NUMA and contention effects in multi-GPU systems,” in GPGPU-4 . ACM, 2011
2011
Earlier work this paper cites.
H. Wang, S. Potluri, M. Luo, A. K. Singh, S. Sur, and D. K. Panda, “MVAPICH2-GPU: optimized GPU to GPU communication for InfiniBand clusters,” Computer Science-Research and Development , 2011
2011
Earlier work this paper cites.
Q. Xu, H. Jeon, and M. Annavaram, “Graph processing on GPU: Where are the bottlenecks?” in International Symposium on Workload Characterization (IISWC) . IEEE, 2014
2014
Earlier work this paper cites.
G. Kim, M. Lee, J. Jeong, and J. Kim, “Multi-GPU system design with memory networks,” in 47th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2014
2014
Earlier work this paper cites.
H. Wang, S. Potluri, D. Bureddy, C. Rosales, and D. K. Panda, “GPU-aware MPI on RDMA-enabled clusters: Design, implementation and evaluation,” IEEE Transactions on Parallel and Distributed Systems , 2014
2014
Earlier work this paper cites.
NVIDIA, “CUDA SDK Code Samples,” 2015
2015
Earlier work this paper cites.
A. Li, Y. Tay, A. Kumar, and H. Corporaal, “Transit: A visual analytical model for multithreaded machines,” in International Symposium on High-Performance Parallel and Distributed Computing (HPDC) . ACM, 2015
2015
Earlier work this paper cites.
A. Li, G.-J. van den Braak, A. Kumar, and H. Corporaal, “Adaptive and transparent cache bypassing for GPUs,” in International Conference for High Performance Computing, Networking, Storage and Analysis (SC) . ACM, 2015
2015
Earlier work this paper cites.
A. Li, G.-J. van den Braak, H. Corporaal, and A. Kumar, “Fine-grained synchronizations and dataflow programming on GPUs,” in International Conference on Supercomputing (ICS) . ACM, 2015
2015
Earlier work this paper cites.
T. Ben-Nun, E. Levy, A. Barak, and E. Rubin, “Memory access patterns: the missing piece of the multi-GPU puzzle,” in International Conference for High Performance Computing, Networking, Storage and Analysis (SC) . ACM, 2015
2015
Earlier work this paper cites.
J. Cabezas, L. Vilanova, I. Gelado, T. B. Jablin, N. Navarro, and W.-m. W. Hwu, “Automatic parallelization of kernels in shared-memory multi-GPU nodes,” in International Conference on Supercomputing (SC) . ACM, 2015
2015
Cited alongside, same era.
A. Li, S. L. Song, E. Brugel, A. Kumar, D. Chavarria-Miranda, and H. Corporaal, “X: A comprehensive analytic model for parallel machines,” in International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 2016
2016
Cited alongside, same era.
A. Li, S. L. Song, M. Wijtvliet, A. Kumar, and H. Corporaal, “SFU-driven transparent approximation acceleration on GPUs,” in International Conference on Supercomputing (ICS) . ACM, 2016
2016
Cited alongside, same era.
A. Li, S. L. Song, A. Kumar, E. Z. Zhang, D. Chavarría-Miranda, and H. Corporaal, “Critical points based register-concurrency autotuning for GPUs,” in Conference on Design, Automation & Test in Europe , 2016
2016
Cited alongside, same era.
A. Li, W. Liu, M. R. Kristensen, B. Vinter, H. Wang, K. Hou, A. Marquez, and S. L. Song, “Exploring and analyzing the real impact of modern on-package memory on HPC scientific kernels,” in International Conference for High Performance Computing, Networking, Storage and Analysis (SC) . ACM, 2017
2017
Later among the works it cites.
A. Li, S. L. Song, W. Liu, X. Liu, A. Kumar, and H. Corporaal, “Locality-aware CTA clustering for modern GPUs,” ACM SIGOPS Operating Systems Review , 2017
2017
Later among the works it cites.
A. Li, W. Zhao, and S. L. Song, “Bvf: enabling significant on-chip power savings via bit-value-favor for throughput processors,” in Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture . ACM, 2017
2017
Later among the works it cites.
B. Klenk, H. Fröening, H. Eberle, and L. Dennison, “Relaxations for high-performance message passing on massively parallel SIMT processors,” in International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Liu, A. Li, J. Hogg, I. S. Duff, and B. Vinter, “A synchronization-free algorithm for parallel sparse triangular solves,” in European Conference on Parallel Processing (EuroPar) . Springer, 2016
2016
Cited alongside, same era.
J. Li, Y. Ma, C. Yan, and R. Vuduc, “Optimizing sparse tensor times matrix on multi-core and many-core architectures,” in Proceedings of the Sixth Workshop on Irregular Applications: Architectures and Algorithms . IEEE Press, 2016
2016
Cited alongside, same era.
T. Gysi, J. Bär, and T. Hoefler, “dCUDA: hardware supported overlap of computation and communication,” in International Conference for High Performance Computing, Networking, Storage and Analysis (SC) . IEEE, 2016
2016
Cited alongside, same era.
2017
Cited alongside, same era.
NVIDIA, “NVIDIA DGX-1 System Architecture White Paper,” 2017
2017
Cited alongside, same era.
“The System Bottleneck Shifts to PCI-Express,” https://www.nextplatform.com/2017/07/14/system-bottleneck-shifts-pci-express/
2017
Cited alongside, same era.
D. Foley and J. Danskin, “Ultra-Performance Pascal GPU and NVLink Interconnect,” IEEE Micro , 2017
2017
Cited alongside, same era.
S. Shams, R. Platania, K. Lee, and S.-J. Park, “Evaluation of Deep Learning Frameworks Over Different HPC Architectures,” in 37th International Conference on Distributed Computing Systems (ICDCS) . IEEE, 2017
2017
Cited alongside, same era.
2017
Later among the works it cites.
2017
Later among the works it cites.
O. Fuhrer, T. Chadha, T. Hoefler, G. Kwasniewski, X. Lapillonne, D. Leutwyler, D. Lüthi, C. Osuna, C. Schär, T. C. Schulthess et al. , “Near-global climate simulation at 1 km resolution: establishing a performance baseline on 4888 gpus with cosmo 5.0,” Geoscientific Model Development , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
NVIDIA, “NVIDIA DGX-2H The World’s Most Powerful System for The Most Complex AI Challenges,” 2018
2018
Later among the works it cites.
A. Li, S. Song, J. Chen, X. Liu, N. Tallent, and K. Barker, “Tartan: Evaluating Modern GPU Interconnect via a Multi-GPU Benchmark Suite,” in International Symposium on Workload Characterization (IISWC) . IEEE, 2018
2018
Later among the works it cites.
E. Agostini, D. Rossetti, and S. Potluri, “GPUDirect Async: Exploring GPU synchronous communication techniques for InfiniBand clusters,” Journal of Parallel and Distributed Computing , 2018
2018
Later among the works it cites.
J. Yin, “Early experiences with Machine Learning and Deep Learning on Summit/SummitDev,” https://www.olcf.ornl.gov/wp-content/uploads/2018/12/summit_training_mldl.pdf
2018
Later among the works it cites.
A. Li, W. Liu, L. Wang, K. Barker, and S. L. Song, “Warp-consolidation: A novel execution model for gpus,” in International Conference on Supercomputing (ICS) . ACM, 2018
2018
Later among the works it cites.
D. Shen, S. L. Song, A. Li, and X. Liu, “Cudaadvisor: Llvm-based runtime profiling for modern gpus,” in Proceedings of the 2018 International Symposium on Code Generation and Optimization . ACM, 2018
2018
Later among the works it cites.
L. Wang, J. Ye, Y. Zhao, W. Wu, A. Li, S. L. Song, Z. Xu, and T. Kraska, “Superneurons: Dynamic GPU memory management for training deep neural networks,” in ACM SIGPLAN Notices . ACM, 2018
2018
Later among the works it cites.
Y. Sun, S. Mukherjee, T. Baruah, S. Dong, J. Gutierrez, P. Mohan, and D. Kaeli, “Evaluating performance tradeoffs on the radeon open compute platform,” in International Symposium on Performance Analysis of Systems and Software (ISPASS) . IEEE, 2018
2018
Later among the works it cites.