Fetching the paper…
Reading the bibliography…
Multi-Instance GPU (MIG) is a new feature introduced by NVIDIA A100 GPUs that partitions one physical GPU into multiple GPU instances.
A genetic algorithm for minimizing the makespan in the case of scheduling identical parallel machines
L. Min and W. Cheng · 1999
Earlier work this paper cites.
Parallel machine scheduling with splitting jobs
W. Xing and J. Zhang · 2000
Earlier work this paper cites.
Solving discrete-continuous scheduling problems by tabu search
J. Józefowska, M. Mika, R. Różycki, G. Waligóra, and J. Węglarz · 2001
Earlier work this paper cites.
Reconfigurable machine tools
R. G. Landers, B.-K. Min, and Y. Koren · 2001
Earlier work this paper cites.
Unrelated parallel machine scheduling with setup times using simulated annealing
D.-W. Kim, K.-H. Kim, W. Jang, and F. F. Chen · 2002
Earlier work this paper cites.
Operating systems for reconfigurable embedded platforms: Online scheduling of real-time tasks
C. Steiger, H. Walder, and M. Platzner · 2004
Earlier work this paper cites.
Z3: An efficient smt solver
L. De Moura and N. Bjørner · 2008
Earlier work this paper cites.
A genetic algorithm for the flexible job-shop scheduling problem
F. Pezzella, G. Morganti, and G. Ciaschetti · 2008
Earlier work this paper cites.
The discrete part of the discrete-continuous scheduling problems-new properties
M. Gorczyca, A. Janiak, and W. Janiak · 2009
Earlier work this paper cites.
Scheduling
M. Pinedo · 2012
Earlier work this paper cites.
Modelling the problem of production scheduling for reconfigurable manufacturing systems
A. Azab and B. Naderi · 2015
Earlier work this paper cites.
Tensorflow-serving: Flexible, high-performance ml serving
C. Olston, N. Fiedel, K. Gorovoy, J. Harmsen, L. Lao, F. Li, V. Rajashekhar, S. Ramesh, and J. Soyke · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Earlier work this paper cites.
https://www.xilinx.com/support/documentation/sw_manuals/xilinx2018_1/ug909-vivado-partial-reconfiguration.pdf , 2018
Vivado Design Suite User Guide Partial Reconfiguration · 2018
Cited alongside, same era.
Low latency rnn inference with cellular batching
P. Gao, L. Yu, Y. Wu, and J. Li · 2018
Cited alongside, same era.
Optimus: an efficient dynamic resource scheduler for deep learning clusters
Y. Peng, Y. Bao, Y. Chen, C. Wu, and C. Guo · 2018
Cited alongside, same era.
A generic communication scheduler for distributed dnn training acceleration
Y. Peng, Y. Zhu, Y. Chen, Y. Bao, B. Yi, C. Lan, C. Wu, and C. Guo · 2019
Cited alongside, same era.
Infaas: A model-less inference serving system
F. Romero, Q. Li, N. J. Yadwadkar, and C. Kozyrakis · 2019
Cited alongside, same era.
https://aws.amazon.com/ec2/instance-types/p3/ , 2021
Amazon EC2 P3 Instances · 2021
Closest in time.
https://aws.amazon.com/ec2/instance-types/p4/ , 2021
Amazon EC2 P4d Instances · 2021
Closest in time.
https://en.wikipedia.org/wiki/Cutting_stock_problem , 2021
Cutting stock problem · 2021
Closest in time.
https://www.nvidia.com/en-us/data-center/gpu-cloud-computing/ , 2021
GPU cloud computing solution · 2021
Closest in time.
https://kubernetes.io/docs/concepts/architecture/controller/ , 2021
Kubernetes Controllers · 2021
Closest in time.
https://developer.nvidia.com/blog/minimizing-dl-inference-latency-with-mig/ , 2021
Minimizing Deep Learning Inference Latency with NVIDIA Multi-Instance GPU · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Shen, L. Chen, Y. Jin, L. Zhao, B. Kong, M. Philipose, A. Krishnamurthy, and R. Sundaram · 2019
Cited alongside, same era.
Balancing efficiency and fairness in heterogeneous gpu clusters for deep learning
S. Chaudhary, R. Ramjee, M. Sivathanu, N. Kwatra, and S. Viswanatha · 2020
Cited alongside, same era.
Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping
C.-C. Huang, G. Jin, and J. Li · 2020
Cited alongside, same era.
Approximation algorithms for scheduling with class constraints
K. Jansen, A. Lassota, and M. Maack · 2020
Cited alongside, same era.
Flexible job shop scheduling problem with reconfigurable machine tools: An improved differential evolution algorithm
M. Mahmoodjanloo, R. Tavakkoli-Moghaddam, A. Baboli, and A. Bozorgi-Amiri · 2020
Cited alongside, same era.
Pollux: Co-adaptive cluster scheduling for goodput-optimized deep learning
A. Qiao, W. Neiswanger, Q. Ho, H. Zhang, G. R. Ganger, and E. P. Xing · 2020
Cited alongside, same era.
Resource partitioning and application scheduling with module merging on dynamically and partially reconfigurable fpgas
Z. Wang, Q. Tang, B. Guo, J.-B. Wei, and L. Wang · 2020
Cited alongside, same era.
https://www.nvidia.com/en-us/data-center/a100/ , 2021
NVIDIA A100 · 2021
Closest in time.
https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/nvidia-ampere-architecture-whitepaper.pdf , 2021
NVIDIA A100 Tensor Core GPU Architecture · 2021
Closest in time.
https://developer.nvidia.com/deep-learning-performance-training-inference , 2021
NVIDIA Data Center Deep Learning Product Performance · 2021
Closest in time.
https://docs.nvidia.com/datacenter/tesla/pdf/NVIDIA_MIG_User_Guide.pdf , 2021
NVIDIA Multi-Instance GPU User Guide · 2021
Closest in time.
https://pytorch.org/hub/ , 2021
PyTorch Hub · 2021
Closest in time.
https://tfhub.dev/ , 2021
TensorFlow Hub · 2021
Closest in time.
Accelerating deep learning inference via learned caches
A. Balasubramanian, A. Kumar, Y. Liu, H. Cao, S. Venkataraman, and A. Akella · 2021
Closest in time.