Fetching the paper…
Reading the bibliography…
GPU remoting is a promising technique for supporting AI applications.
An efficient implementation of GPU virtualization in high performance clusters
2009
Earlier work this paper cites.
Gvim: Gpu-accelerated virtual machines
2009
Earlier work this paper cites.
Modeling the cuda remoting virtualization behaviour in high performance networks
2010
Earlier work this paper cites.
rcuda: Reducing the number of gpu-based accelerators in high performance clusters
2010
Earlier work this paper cites.
A GPGPU transparent virtualization component for high performance computing clouds
2010
Earlier work this paper cites.
Performance of CUDA virtualized remote gpus in high performance clusters
2011
Earlier work this paper cites.
Pegasus: Coordinated scheduling for virtualized accelerator-based systems
2011
Earlier work this paper cites.
A virtual memory based runtime to support multi-tenancy in clusters with gpus
2012
Earlier work this paper cites.
vcuda: Gpu-accelerated high-performance computing in virtual machines
2012
Earlier work this paper cites.
FaRM: Fast remote memory
2014
Earlier work this paper cites.
Mica: A holistic approach to fast in-memory key-value storage
2014
Earlier work this paper cites.
vgasa: Adaptive scheduling algorithm of virtualized GPU resource in cloud gaming
2014
Earlier work this paper cites.
Exploring the suitability of remote GPGPU virtualization for the openacc programming model using rcuda
2015
Earlier work this paper cites.
Network requirements for resource disaggregation
2016
Earlier work this paper cites.
RDMA over commodity ethernet at scale
2016
Earlier work this paper cites.
Design guidelines for high performance RDMA systems
2016
Earlier work this paper cites.
GPU virtualization and scheduling methods: A comprehensive survey
2017
Cited alongside, same era.
Lite kernel rdma support for datacenter applications
2017
Cited alongside, same era.
Ray: A distributed framework for emerging AI applications
2018
Cited alongside, same era.
Legoos: A disseminated, distributed OS for hardware resource disaggregation
2018
Cited alongside, same era.
Cachecloud: Towards speed-of-light datacenter communication
2018
Cited alongside, same era.
Deconstructing RDMA-enabled distributed transactions: Hybrid is better!
2018
Cited alongside, same era.
ConnectX-7 product brief
2022
Later among the works it cites.
Singularity: Planet-scale, preemptive and elastic scheduling of AI workloads
2022
Later among the works it cites.
The ai community building the future
2023
Later among the works it cites.
Dxpu: Large scale disaggregated GPU pools in the datacenter
2023
Later among the works it cites.
Skadi: Building a distributed runtime for data systems in disaggregated data centers
2023
Later among the works it cites.
AlpaServe: Statistical multiplexing with model parallelism for deep learning serving
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
AIFM: high-performance, application-integrated far memory
2020
Cited alongside, same era.
Disaggregating persistent memory and controlling them remotely: An exploration of passive disaggregated key-value stores
2020
Cited alongside, same era.
Semeru: A memory-disaggregated managed runtime
2020
Cited alongside, same era.
Polardb serverless: A cloud native database for disaggregated data centers
2021
Cited alongside, same era.
Efficient large-scale language model training on gpu clusters using megatron-lm
2021
Cited alongside, same era.
Paella: Low-latency model serving with software-defined GPU scheduling
2023
Later among the works it cites.
Efficiently scaling transformer inference
2023
Later among the works it cites.
Mware vsphere bitfusion
2023
Later among the works it cites.
Faaswap: Slo-aware, gpu-efficient serverless inference via model swapping
2023
Later among the works it cites.
Partial failure resilient memory management system for (cxl-based) distributed shared memory
2023
Later among the works it cites.
NVIDIA DGX Platform
2024
Closest in time.
Open Fabrics Enterprise Distribution (OFED) Performance Tests
2024
Closest in time.
Datasets and dataloaders
2024
Closest in time.
Torchelastic
2024
Closest in time.