Fetching the paper…
Reading the bibliography…
Cloud platforms today have been deploying hardware accelerators like neural processing units (NPUs) for powering machine learning (ML) inference services.
P. Barham, B. Dragovic, K. Fraser, S. Hand, T. Harris, A. Ho, R. Neugebauer, I. Pratt, and A. Warfield, “Xen and the art of virtualization,” in Proceedings of the Nineteenth ACM Symposium on Operating Systems Principles (SOSP’03) , Bolton Landing, NY, USA, 2003
2003
Earlier work this paper cites.
K. Adams and O. Agesen, “A comparison of software and hardware techniques for x86 virtualization,” in Proceedings of the 12th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS’06) , San Jose, CA, USA, 2006
2006
Earlier work this paper cites.
R. Bhargava, B. Serebrin, F. Spadini, and S. Manne, “Accelerating two-dimensional page walks for virtualized systems,” in Proceedings of the 13th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS’08) , Seattle, WA, USA, 2008
2008
Earlier work this paper cites.
VMWare, “Performance Evaluation of Intel EPT Hardware Assist,” 2009. [Online]. Available: https://www.vmware.com/pdf/Perf_ESX_Intel-EPT-eval.pdf
2009
Earlier work this paper cites.
VMWare, “vSphere Networking,” 2009. [Online]. Available: https://docs.vmware.com/en/VMware-vSphere/8.0/vsphere-esxi-vcenter-802-networking-guide.pdf
2009
Earlier work this paper cites.
J. Mars, L. Tang, R. Hundt, K. Skadron, and M. L. Soffa, “Bubble-up: Increasing utilization in modern warehouse scale computers via sensible co-locations,” in Proceedings of the 44th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO’11) , Porto Alegre, Brazil, 2011
2011
Earlier work this paper cites.
T. Chen, Z. Du, N. Sun, J. Wang, C. Wu, Y. Chen, and O. Temam, “DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning,” in Proceedings of the 20th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS’14) , Salt Lake City, UT, 2014
2014
Earlier work this paper cites.
T. P. P. de Lacerda Ruivo, G. B. Altayo, G. Garzoglio, S. Timm, H. W. Kim, S.-Y. Noh, and I. Raicu, “Exploring infiniband hardware virtualization in opennebula towards efficient high-performance computing,” in Proceedings of the 14th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid’14) , Chicago, IL, USA, 2014
2014
Earlier work this paper cites.
I. Tanasic, I. Gelado, J. Cabezas, A. Ramirez, N. Navarro, and M. Valero, “Enabling preemptive multiprogramming on gpus,” in 2014 ACM/IEEE 41st International Symposium on Computer Architecture (ISCA) , 2014
2014
Earlier work this paper cites.
D. Lo, L. Cheng, R. Govindaraju, P. Ranganathan, and C. Kozyrakis, “Heracles: Improving resource efficiency at scale,” in Proceedings of the 42nd Annual International Symposium on Computer Architecture (ISCA’15) , Portland, OR, USA, 2015
2015
Earlier work this paper cites.
Q. Chen, H. Yang, J. Mars, and L. Tang, “Baymax: Qos awareness and increased utilization for non-preemptive accelerators in warehouse scale computers,” in Proceedings of the Twenty-First International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS’16) , Atlanta, GA, 2016
2016
Earlier work this paper cites.
Z. Lin, L. Nyland, and H. Zhou, “Enabling efficient preemption for simt architectures with lightweight context switching,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC’16) , Salt Lake City, UT, USA, 2016
2016
Earlier work this paper cites.
Q. Chen, H. Yang, M. Guo, R. S. Kannan, J. Mars, and L. Tang, “Prophet: Precise qos prediction on non-preemptive accelerators to improve utilization in warehouse-scale computers,” in Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS’17) , Xi’an, China, 2017
2017
Earlier work this paper cites.
E. Chung, J. Fowers, K. Ovtcharov, M. Papamichael, A. Caulfield, T. Massengil, M. Liu, D. Lo, S. Alkalay, M. Haselman, C. Boehn, O. Firestein, A. Forin, K. S. Gatlin, M. Ghandi, S. Heil, K. Holohan, T. Juhasz, R. K. Kovvuri, S. Lanka, F. van Megen, D. Mukhortov, P. Patel, S. Reinhardt, A. Sapek, R. Seera, B. Sridharan, L. Woods, P. Yi-Xiao, R. Zhao, and D. Burger, “Accelerating Persistent Neural Networks at Datacenter Scale,” in Proceedings of HotChips’17 , Cupertino, CA, USA, 2017
2017
Earlier work this paper cites.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P.-l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V. Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D. H. Yoon, “In-Datacenter Performance Analysis of a Tensor Processing Unit,” in Proceedings of the 44th International Symposium on Computer Architecture (ISCA’17) , Toronto, Canada, 2017
2017
Earlier work this paper cites.
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic Differentiation in PyTorch,” in Proceedings of the 30th International Conference on Neural Information Processing Systems (NIPS’17) , Long Beach, CA, USA, 2017
2017
Earlier work this paper cites.
T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, H. Shen, M. Cowan, L. Wang, Y. Hu, L. Ceze, C. Guestrin, and A. Krishnamurthy, “TVM: An Automated End-to-End Optimizing Compiler for Deep Learning,” in Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI’18) , Carlsbad, CA, 2018
2018
Earlier work this paper cites.
A. Khawaja, J. Landgraf, R. Prakash, M. Wei, E. Schkufza, and C. J. Rossbach, “Sharing, protection, and compatibility for reconfigurable fabric with AmorphOS,” in 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI’18) , Carlsbad, CA, USA, 2018
2018
Earlier work this paper cites.
H. Liao, J. Tu, J. Xia, and X. Zhou, “Davinci: A scalable architecture for neural network computing,” in 2019 IEEE Hot Chips 31 Symposium (HCS) , Los Alamitos, CA, USA, 2019
2019
Earlier work this paper cites.
E. Baek, D. Kwon, and J. Kim, “A multi-neural network acceleration architecture,” in Proceedings of the ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA’20) , Virtual Event, 2020
2020
Earlier work this paper cites.
Y. Choi and M. Rhu, “PREMA: A predictive multi-task scheduling algorithm for preemptible neural processing units,” in Proceedings of the 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA’20) , San Diego, CA, USA, 2020
2020
Earlier work this paper cites.
S. Ghodrati, B. H. Ahn, J. Kyung Kim, S. Kinzer, B. R. Yatham, N. Alla, H. Sharma, M. Alian, E. Ebrahimi, N. S. Kim, C. Young, and H. Esmaeilzadeh, “Planaria: Dynamic architecture fission for spatial multi-tenant acceleration of deep neural networks,” in Proceedings of the 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO’20) , Virtual Event, 2020
2020
Cited alongside, same era.
L. Gwennap, “Tenstorrent scales ai performance: New multicore architecture leads in data-center power efficiency,” 2020. [Online]. Available: https://www.linleygroup.com/mpr/article.php?id=12287
2020
Cited alongside, same era.
C.-C. Huang, G. Jin, and J. Li, “Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping,” in Proceedings of the 25th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS’20) , Lausanne, Switzerland, 2020
2020
Cited alongside, same era.
Nvidia, “Multi-Instance GPU user guide,” 2022. [Online]. Available: https://docs.nvidia.com/datacenter/tesla/mig-user-guide/
2022
Later among the works it cites.
Nvidia, “Virtual GPU software user guide,” 2022. [Online]. Available: https://docs.nvidia.com/grid/latest/grid-vgpu-user-guide/
2022
Later among the works it cites.
E. Onose, “Machine learning as a service: What it is, when to use it and what are the best tools out there,” 2022. [Online]. Available: https://neptune.ai/blog/machine-learning-as-a-service-what-it-is-when-to-use-it-and-what-are-the-best-tools-out-there
2022
Later among the works it cites.
RUN:AI, “Google TPU Architecture and Performance Best Practices,” 2022. [Online]. Available: https://www.run.ai/guides/cloud-deep-learning/google-tpu
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Hui, “AI Chips Technology Trends and Landscapes (Mobile SoC, Intel, Asian AI Chips, Low-Power Inference Chips),” 2020. [Online]. Available: https://jonathan-hui.medium.com/ai-chips-technology-trends-landscape-mobile-soc-intel-asian-ai-chips-low-power-inference-4db701dbe85d
2020
Cited alongside, same era.
Y. Jiao, L. Han, and X. Long, “Hanguang 800 npu – the ultimate ai inference solution for data centers,” in 2020 IEEE Hot Chips 32 Symposium (HCS) , Palo Alto, CA, USA, 2020
2020
Cited alongside, same era.
N. P. Jouppi, D. H. Yoon, G. Kurian, S. Li, N. Patil, J. Laudon, C. Young, and D. Patterson, “A domain-specific supercomputer for training deep neural networks,” Commun. ACM , vol. 63, no. 7, June 2020
2020
Cited alongside, same era.
J. Ma, G. Zuo, K. Loughlin, X. Cheng, Y. Liu, A. M. Eneyew, Z. Qi, and B. Kasikci, “A hypervisor for shared-memory fpga platforms,” in Proceedings of the 25th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS’20) , Lausanne, Switzerland, 2020
2020
Cited alongside, same era.
V. J. Reddi, C. Cheng, D. Kanter, P. Mattson, G. Schmuelling, C.-J. Wu, B. Anderson, M. Breughe, M. Charlebois, W. Chou, R. Chukka, C. Coleman, S. Davis, P. Deng, G. Diamos, J. Duke, D. Fick, J. S. Gardner, I. Hubara, S. Idgunji, T. B. Jablin, J. Jiao, T. S. John, P. Kanwar, D. Lee, J. Liao, A. Lokhmotov, F. Massa, P. Meng, P. Micikevicius, C. Osborne, G. Pekhimenko, A. T. R. Rajan, D. Sequeira, A. Sirasao, F. Sun, H. Tang, M. Thomson, F. Wei, E. Wu, L. Xu, K. Yamada, B. Yu, G. Yuan, A. Zhong, P. Zhang, and Y. Zhou, “Mlperf inference benchmark,” 2020
2020
Cited alongside, same era.
H. Yu, A. M. Peters, A. Akshintala, and C. J. Rossbach, “AvA: Accelerated virtualization of accelerators,” in Proceedings of the 25th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS’20) , Lausanne, Switzerland, 2020
2020
Cited alongside, same era.
Y. Zha and J. Li, “Virtualizing fpgas in the cloud,” in Proceedings of the 25th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS’20) , Lausanne, Switzerland, 2020
2020
Cited alongside, same era.
Altexsoft, “Comparing Machine Learning as a Service: Amazon, Microsoft Azure, Google Cloud AI, IBM Watson,” 2021. [Online]. Available: https://www.altexsoft.com/blog/datascience/comparing-machine-learning-as-a-service-amazon-microsoft-azure-google-cloud-ai-ibm-watson/
2021
Cited alongside, same era.
J. Landgraf, T. Yang, W. Lin, C. J. Rossbach, and E. Schkufza, “Compiler-driven fpga virtualization with synergy,” in Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS’21) , Virtual Event, 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
K. Wiggers, “Microsoft and nvidia team up to build new azure-hosted ai supercomputer,” 2022. [Online]. Available: https://techcrunch.com/2022/11/16/microsoft-and-nvidia-team-up-to-build-new-azure-hosted-ai-supercomputer/
2022
Later among the works it cites.
H. Zhu, R. Wu, Y. Diao, S. Ke, H. Li, C. Zhang, J. Xue, L. Ma, Y. Xia, W. Cui, F. Yang, M. Yang, L. Zhou, A. Cidon, and G. Pekhimenko, “ROLLER: Fast and efficient tensor compilation for deep learning,” in Proceedings of 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI’22) , Carlsbad, CA, USA, 2022
2022
Later among the works it cites.
AMD, “AI Engine: Meeting the Compute Demands of Next-Generation Applications,” 2023. [Online]. Available: https://www.xilinx.com/products/technology/ai-engine.html
2023
Later among the works it cites.
A. AWS, “Aws inferentia,” 2023. [Online]. Available: https://aws.amazon.com/machine-learning/inferentia/
2023
Later among the works it cites.
Google, “Create production-grade machine learning models with TensorFlow,” 2023. [Online]. Available: https://www.tensorflow.org/
2023
Later among the works it cites.
Google, “Supported reference models,” 2023. [Online]. Available: https://cloud.google.com/tpu/docs/tutorials/supported-models
2023
Later among the works it cites.
Google, “XLA: Optimizing Compiler for Machine Learning,” 2023. [Online]. Available: https://www.tensorflow.org/xla
2023
Later among the works it cites.
N. Jia and K. Wankhede, “Vfio mediated devices,” 2023. [Online]. Available: https://docs.kernel.org/driver-api/vfio-mediated-device.html
2023
Later among the works it cites.
The KubeVirt Contributors, “Kubevirt.io,” 2023. [Online]. Available: https://kubevirt.io/
2023
Later among the works it cites.
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom, “Llama 2: Open foundation and fine-tuned chat models,” 2023
2023
Later among the works it cites.
A. Vahdat and M. Lohmeyer, “Enabling next-generation ai workloads: Announcing tpu v5p and ai hypercomputer,” 2023. [Online]. Available: https://cloud.google.com/blog/products/ai-machine-learning/introducing-cloud-tpu-v5p-and-ai-hypercomputer
2023
Later among the works it cites.
Wolfram Alpha LLC, “Wolframalpha: Computational intelligence,” 2023. [Online]. Available: https://www.wolframalpha.com/
2023
Later among the works it cites.
Y. Xue, Y. Liu, and J. Huang, “System virtualization for neural processing units,” in Proceedings of the 19th Workshop on Hot Topics in Operating Systems (HotOS’23) , Providence, RI, USA, 2023
2023
Later among the works it cites.
Y. Xue, Y. Liu, L. Nai, and J. Huang, “Hardware-assisted virtualization for neural processing units,” in The 1st Workshop on Hot Topics in System Infrastructure (HotInfra’23) , Orlando, FL, USA, 2023
2023
Later among the works it cites.
Y. Xue, Y. Liu, L. Nai, and J. Huang, “V10: Hardware-assisted npu multi-tenancy for improved resource utilization and fairness,” in Proceedings of the 50th Annual International Symposium on Computer Architecture (ISCA’23) , Orlando, FL, USA, 2023
2023
Later among the works it cites.
H. Zhang, Y. Zhou, Y. Xue, Y. Liu, and J. Huang, “G10: Enabling an efficient unified gpu memory and storage architecture with smart tensor migrations,” in Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO’23) , Toronto, ON, Canada, 2023
2023
Later among the works it cites.