Fetching the paper…
Reading the bibliography…
Heterogeneous supercomputers have become the standard in HPC.
T. Brecht, “On the importance of parallel application placement in numa multiprocessors,” in Symposium on Experiences with Distributed and Multiprocessor Systems (SEDMS IV)
1993
Earlier work this paper cites.
J. D. McCalpin et al
1995
Earlier work this paper cites.
J. Kim, W. J. Dally, S. Scott, and D. Abts, “Technology-driven, highly-scalable dragonfly topology,” ACM SIGARCH Computer Architecture News
2008
Earlier work this paper cites.
S. Matsuoka, T. Aoki, T. Endo, A. Nukada, T. Kato, and A. Hasegawa, “Gpu accelerated computing–from hype to mainstream, the rebirth of vector computing,” in Journal of Physics: Conference Series
2009
Earlier work this paper cites.
S. Che, M. Boyer, J. Meng, D. Tarjan, J. W. Sheaffer, S.-H. Lee, and K. Skadron, “Rodinia: A benchmark suite for heterogeneous computing,” in 2009 IEEE international symposium on workload characterization (IISWC)
2009
Earlier work this paper cites.
A. Danalis, G. Marin, C. McCurdy, J. S. Meredith, P. C. Roth, K. Spafford, V. Tipparaju, and J. S. Vetter, “The scalable heterogeneous computing (shoc) benchmark suite,” in Proceedings of the 3rd workshop on general-purpose computation on graphics processing units
2010
Earlier work this paper cites.
M. Bauer, H. Cook, and B. Khailany, “Cudadma: optimizing gpu memory bandwidth via warp specialization,” in Proceedings of 2011 international conference for high performance computing, networking, storage and analysis
2011
Earlier work this paper cites.
J. A. Stratton, C. Rodrigues, I.-J. Sung, N. Obeid, L.-W. Chang, N. Anssari, G. D. Liu, and W.-m. W. Hwu, “Parboil: A revised benchmark suite for scientific and commercial throughput computing,” Center for Reliable and High-Performance Computing
2012
Earlier work this paper cites.
S. Kato, J. Aumiller, and S. Brandt, “Zero-copy i/o processing for low-latency gpu computing,” in Proceedings of the ACM/IEEE 4th International Conference on Cyber-Physical Systems
2013
Earlier work this paper cites.
J. Shen, A. L. Varbanescu, H. Sips, M. Arntzen, and D. G. Simons, “Glinda: A framework for accelerating imbalanced applications on heterogeneous platforms,” in Proceedings of the ACM International Conference on Computing Frontiers
2013
Earlier work this paper cites.
P. Mistry, Y. Ukidave, D. Schaa, and D. Kaeli, “Valar: A benchmark suite to study the dynamic behavior of heterogeneous systems,” in Proceedings of the 6th Workshop on General Purpose Processor Using Graphics Processing Units
2013
Earlier work this paper cites.
C. A. Navarro, N. Hitschfeld-Kahler, and L. Mateu, “A survey on parallel computing and its applications in data-parallel problems using gpu architectures,” Communications in Computational Physics
2014
Earlier work this paper cites.
B. Van Werkhoven, J. Maassen, F. J. Seinstra, and H. E. Bal, “Performance models for cpu-gpu data transfers,” in 2014 14th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing
2014
Earlier work this paper cites.
J. Shen, A. L. Varbanescu, and H. Sips, “Look before you leap: Using the right hardware resources to accelerate applications,” in 2014 IEEE Intl Conf on High Performance Computing and Communications, 2014 IEEE 6th Intl Symp on Cyberspace Safety and Security, 2014 IEEE 11th Intl Conf on Embedded Software and Syst (HPCC, CSS, ICESS)
2014
Earlier work this paper cites.
S. Mittal and J. S. Vetter, “A survey of cpu-gpu heterogeneous computing techniques,” ACM Computing Surveys (CSUR)
2015
Cited alongside, same era.
J. Shen, A. L. Varbanescu, Y. Lu, P. Zou, and H. Sips, “Workload partitioning for accelerating applications on heterogeneous platforms,” IEEE Transactions on Parallel and Distributed Systems
2015
Cited alongside, same era.
D. Unat, A. Dubey, T. Hoefler, J. Shalf, M. Abraham, M. Bianco, B. L. Chamberlain, R. Cledat, H. C. Edwards, H. Finkel, et al
2017
Cited alongside, same era.
S. Ramos and T. Hoefler, “Capability models for manycore memory systems: A case-study with xeon phi knl,” in 2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
2017
Cited alongside, same era.
U. Milic, O. Villa, E. Bolotin, A. Arunkumar, E. Ebrahimi, A. Jaleel, A. Ramirez, and D. Nellans, “Beyond the socket: Numa-aware gpus,” in Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture
C. Pearson, A. Dakkak, S. Hashash, C. Li, I.-H. Chung, J. Xiong, and W.-M. Hwu, “Evaluating characteristics of cuda communication primitives on high-bandwidth interconnects,” in Proceedings of the 2019 ACM/SPEC International Conference on Performance Engineering
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al
2019
Later among the works it cites.
D. De Sensi, S. Di Girolamo, K. H. McMahon, D. Roweth, and T. Hoefler, “An in-depth analysis of the slingshot interconnect,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis
2020
Later among the works it cites.
A. Ivanov, N. Dryden, T. Ben-Nun, S. Li, and T. Hoefler, “Data movement is all you need: A case study on optimizing transformers,” Proceedings of Machine Learning and Systems
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
M. Dashti and A. Fedorova, “Analyzing memory management methods on integrated cpu-gpu systems,” in Proceedings of the 2017 ACM SIGPLAN International Symposium on Memory Management
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems
2017
Cited alongside, same era.
L. Wang, J. Ye, Y. Zhao, W. Wu, A. Li, S. L. Song, Z. Xu, and T. Kraska, “Superneurons: Dynamic gpu memory management for training deep neural networks,” in Proceedings of the 23rd ACM SIGPLAN symposium on principles and practice of parallel programming
2018
Cited alongside, same era.
S. Haria, M. D. Hill, and M. M. Swift, “Devirtualizing memory in heterogeneous systems,” in Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and Operating Systems
2018
Cited alongside, same era.
S. Markidis, S. W. Der Chien, E. Laure, I. B. Peng, and J. S. Vetter, “Nvidia tensor core programmability, performance & precision,” in 2018 IEEE international parallel and distributed processing symposium workshops (IPDPSW)
2018
Cited alongside, same era.
A. Li, S. L. Song, J. Chen, J. Li, X. Liu, N. R. Tallent, and K. J. Barker, “Evaluating modern gpu interconnect: Pcie, nvlink, nv-sli, nvswitch and gpudirect,” IEEE Transactions on Parallel and Distributed Systems
2019
Cited alongside, same era.
T. Ben-Nun, J. de Fine Licht, A. N. Ziogas, T. Schneider, and T. Hoefler, “Stateful dataflow multigraphs: A data-centric model for performance portability on heterogeneous architectures,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
2019
Cited alongside, same era.
T. Hoefler, D. Alistarh, T. Ben-Nun, N. Dryden, and A. Peste, “Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks,” Journal of Machine Learning Research
2021
Later among the works it cites.
M. A. Giorgetta, W. Sawyer, X. Lapillonne, P. Adamidis, D. Alexeev, V. Clément, R. Dietlicher, J. F. Engels, M. Esch, H. Franke, et al
2022
Later among the works it cites.
2022
Later among the works it cites.
M. Isaev, N. McDonald, and R. Vuduc, “Scaling infrastructure to support multi-trillion parameter llm training,” in Architecture and System Support for Transformer Models (ASSYST@ ISCA 2023)
2023
Later among the works it cites.
L. Zhang, M. Wahib, P. Chen, J. Meng, X. Wang, T. Endo, and S. Matsuoka, “Perks: a locality-optimized execution model for iterative memory-bound gpu applications,” in Proceedings of the 37th International Conference on Supercomputing
2023
Later among the works it cites.
Y. Wei, Y. C. Huang, H. Tang, N. Sankaran, I. Chadha, D. Dai, O. Oluwole, V. Balan, and E. Lee, “9.3 nvlink-c2c: A coherent off package chip-to-chip interconnect with 40gbps/pin single-ended signaling,” in 2023 IEEE International Solid-State Circuits Conference (ISSCC)
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Li, Y. Wang, X. Liang, and H. Liu, “Automatic blas offloading on unified memory architecture: A study on nvidia grace-hopper,” in Practice and Experience in Advanced Research Computing 2024: Human Powered Computing
2024
Closest in time.