Fetching the paper…
Reading the bibliography…
Memory management across discrete CPU and GPU physical memory is traditionally achieved through explicit GPU allocations and data copy or unified virtual memory.
Rodinia: A benchmark suite for heterogeneous computing. In 2009 IEEE international symposium on workload characterization (IISWC) . IEEE, 44–54
Shuai Che, Michael Boyer, Jiayuan Meng, David Tarjan, Jeremy W Sheaffer, Sang-Ha Lee, and Kevin Skadron. 2009 · 2009
Earlier work this paper cites.
Exploring application performance on emerging hybrid-memory supercomputers. In 2016 IEEE 18th International Conference on High Performance Computing and Communications; IEEE 14th International Conference on Smart City; IEEE 2nd International Conference on Data Science and Systems . IEEE, 473–480
Ivy Bo Peng, Stefano Markidis, Erwin Laure, Gokcen Kestor, and Roberto Gioiosa. 2016 · 2016
Earlier work this paper cites.
Thermostat: Application-transparent page management for two-tiered main memory. In Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems . 631–644
Neha Agarwal and Thomas F Wenisch. 2017 · 2017
Earlier work this paper cites.
Qiskit: An Open-Source Framework for Quantum Computing
Qiskit Community. 2017 · 2017
Earlier work this paper cites.
Siena: Exploring the design space of heterogeneous memory systems. In SC18: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 427–440
Ivy B Peng and Jeffrey S Vetter. 2018 · 2018
Earlier work this paper cites.
Performance evaluation of advanced features in CUDA unified memory. In 2019 IEEE/ACM Workshop on Memory Centric High Performance Computing (MCHPC) . IEEE, 50–57
Steven Chien, Ivy Peng, and Stefano Markidis. 2019 · 2019
Earlier work this paper cites.
Interplay between hardware prefetcher and page eviction policy in cpu-gpu unified virtual memory. In Proceedings of the 46th International Symposium on Computer Architecture . 224–235
Debashis Ganguly, Ziyu Zhang, Jun Yang, and Rami Melhem. 2019 · 2019
Earlier work this paper cites.
Comparing managed memory and ats with and without prefetching on nvidia volta gpus. In 2019 IEEE/ACM Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS) . IEEE, 41–46
Rahulkumar Gayatri, Kevin Gott, and Jack Deslippe. 2019 · 2019
Earlier work this paper cites.
Unified Memory
Nvidia. 2019 · 2019
Earlier work this paper cites.
CUDA Unified Memory
ORNL. 2019 · 2019
Earlier work this paper cites.
Evaluating characteristics of CUDA communication primitives on high-bandwidth interconnects. In Proceedings of the 2019 ACM/SPEC International Conference on Performance Engineering . 209–218
Carl Pearson, Abdul Dakkak, Sarah Hashash, Cheng Li, I-Hsin Chung, Jinjun Xiong, and Wen-Mei Hwu. 2019 · 2019
Earlier work this paper cites.
Nimble page management for tiered memory systems. In Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems . 331–345
Zi Yan, Daniel Lustig, David Nellans, and Abhishek Bhattacharjee. 2019 · 2019
Cited alongside, same era.
Atmem: Adaptive data placement in graph applications on heterogeneous memories. In Proceedings of the 18th ACM/IEEE International Symposium on Code Generation and Optimization . 293–304
Yu Chen, Ivy Peng, Zhen Peng, Xu Liu, and Bin Ren. 2020 · 2020
Cited alongside, same era.
Adaptive page migration for irregular data-intensive applications under gpu memory oversubscription. In 2020 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 451–461
Debashis Ganguly, Ziyu Zhang, Jun Yang, and Rami Melhem. 2020 · 2020
Cited alongside, same era.
Batch-aware unified memory management in GPUs for irregular workloads. In Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems . 1357–1370
Hyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi, and Hyesoon Kim. 2020 · 2020
Early-Adaptor: An Adaptive Framework for Proactive UVM Memory Management. In 2023 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) . IEEE, 248–258
Seokjin Go, Hyunwuk Lee, Junsung Kim, Jiwon Lee, Myung Kuk Yoon, and Won Woo Ro. 2023 · 2023
Later among the works it cites.
Deep learning based data prefetching in CPU-GPU unified virtual memory
Xinjian Long, Xiangyang Gong, Bo Zhang, and Huiyang Zhou. 2023 · 2023
Later among the works it cites.
Tpp: Transparent page placement for cxl-enabled tiered-memory. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 . 742–755
Hasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner, Niket Agarwal, Pallab Bhattacharya, Chris Petersen, Mosharaf Chowdhury, Shobhit Kanaujia, and Prakash Chauhan. 2023 · 2023
Later among the works it cites.
A quantitative approach for adopting disaggregated memory in HPC systems. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . 1–14
Jacob Wahlgren, Gabin Schieffer, Maya Gokhale, and Ivy Peng. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pump up the volume: Processing large data on gpus with fast interconnects. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data . 1633–1649
Clemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl, and Volker Markl. 2020 · 2020
Cited alongside, same era.
Demystifying gpu uvm cost with deep runtime and workload analysis. In 2021 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 141–150
Tyler Allen and Rong Ge. 2021 · 2021
Cited alongside, same era.
Cori: Dancing to the right beat of periodic data movements over hybrid memory systems. In 2021 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 350–359
Thaleia Dimitra Doudali, Daniel Zahka, and Ada Gavrilovska. 2021 · 2021
Cited alongside, same era.
Sentinel: Efficient tensor migration and allocation on heterogeneous memory systems for deep learning. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 598–611
Jie Ren, Jiaolin Luo, Kai Wu, Minjia Zhang, Hyeran Jeon, and Dong Li. 2021b · 2021
Cited alongside, same era.
MD-HM: memoization-based molecular dynamics simulations on big memory system. In Proceedings of the ACM International Conference on Supercomputing . 215–226
Zhen Xie, Wenqian Dong, Jie Liu, Ivy Peng, Yanbao Ma, and Dong Li. 2021 · 2021
Cited alongside, same era.
Evaluating emerging CXL-enabled memory pooling for HPC systems. In 2022 IEEE/ACM Workshop on Memory Centric High Performance Computing (MCHPC) . IEEE, 11–20
Jacob Wahlgren, Maya Gokhale, and Ivy B Peng. 2022 · 2022
Cited alongside, same era.
Quantum Computer Simulations at Warp Speed: Assessing the Impact of GPU Acceleration: A Case Study with IBM Qiskit Aer, Nvidia Thrust & cuQuantum. In 2023 IEEE 19th International Conference on e-Science (e-Science) . 1–10
Jennifer Faj, Ivy Peng, Jacob Wahlgren, and Stefano Markidis. 2023 · 2023
Cited alongside, same era.
Optimizing large-scale plasma simulations on persistent memory-based heterogeneous memory with effective data placement across memory hierarchy. In Proceedings of the ACM International Conference on Supercomputing . 203–214
Jie Ren, Jiaolin Luo, Ivy Peng, Kai Wu, and Dong Li. 2021a
Cited in the paper.
NVIDIA Grace Superchip Early Evaluation for HPC Applications. In Proceedings of the International Conference on High Performance Computing in Asia-Pacific Region Workshops . 45–54
Fabio Banchelli, Joan Vinyals-Ylla-Catala, Josep Pocurull, Marc Clascà, Kilian Peiro, Filippo Spiga, Marta Garcia-Gasulla, and Filippo Mantovani. 2024 · 2024
Closest in time.
CachedArrays: Optimizing Data Movement for Heterogeneous Memory Systems. 38th IEEE International Parallel and Distributed Processing Symposium (IPDPS)
Mark Hildebrand, Jason Lowe-Power, and Venkatesh Akella. 2024 · 2024
Closest in time.
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper. In Practice and Experience in Advanced Research Computing (PEARC’24)
Junjie Li, Yinzhi Wang, Xiao Liang, and Hang Liu. 2024 · 2024
Closest in time.
NVIDIA Grace Hopper Superchip Architecture Whitepaper
Nvidia. 2024a · 2024
Closest in time.
NVIDIA Grace Performance Tuning Guide
Nvidia. 2024b · 2024
Closest in time.
First Impressions of the NVIDIA Grace CPU Superchip and NVIDIA Grace Hopper Superchip for Scientific Workloads. In Proceedings of the International Conference on High Performance Computing in Asia-Pacific Region Workshops . 36–44
Nikolay A Simakov, Matthew D Jones, Thomas R Furlani, Eva Siegmann, and Robert J Harrison. 2024 · 2024
Closest in time.