Fetching the paper…
Reading the bibliography…
To amortize cost, cloud vendors providing DNN acceleration as a service to end-users employ consolidation and virtualization to share the underlying resources among multiple DNN service requests.
J. E. Smith and A. R. Pleszkun, “Implementing Precise Interrupts in Pipelined Processors,” IEEE Transactions on computers
1988
Earlier work this paper cites.
S. Eyerman and L. Eeckhout, “System-level performance metrics for multiprogram workloads,” IEEE micro
2008
Earlier work this paper cites.
NVIDIA, “cuBLAS Library,” 2008
2008
Earlier work this paper cites.
L. Wang, M. Huang, and T. El-Ghazawi, “Exploiting Concurrent Kernel Execution on Graphic Processing Units,” in HPCS
2011
Earlier work this paper cites.
P. Rosenfeld, E. Cooper-Balis, and B. Jacob, “DRAMSim2: A Cycle Accurate Memory System Simulator,” 2011
2011
Earlier work this paper cites.
C. Gregg, J. Dorn, K. Hazelwood, and K. Skadron, “Fine-grained Resource Sharing for Concurrent GPGPU Kernels,” in HotPar
2012
Earlier work this paper cites.
N. Chatterjee, R. Balasubramonian, M. Shevgoor, S. Pugsley, A. Udipi, A. Shafiee, K. Sudan, M. Awasthi, and Z. Chishti, “USIMM: the Utah SImulated Memory Module,” 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” in Proceedings of the International Conference on Neural Information Processing Systems (NIPS)
2012
Earlier work this paper cites.
K. Liu, W. Li, and M. Guo, “Emoticon Smoothed Language Models for Twitter Sentiment Analysis,” in Proceedings of the AAAI Conference on Artificial Intelligence
2012
Earlier work this paper cites.
S. Pai, M. J. Thazhuthaveetil, and R. Govindarajan, “Improving GPGPU Concurrency with Elastic Kernels,” in Proceedings of the International Conference on Architectural Support for Programming Languages and Operation Systems (ASPLOS)
2013
Earlier work this paper cites.
S. Chetlur, C. Woolley, P. Vandermersch, J. Cohen, J. Tran, B. Catanzaro, and E. Shelhamer, “cuDNN: Efficient Primitives for Deep Learning,” in Proceedings of the International Conference on Neural Information Processing Systems (NIPS)
2014
Earlier work this paper cites.
I. Tanasic, I. Gelado, J. Cabezas, A. Ramirez, N. Navarro, and M. Valero, “Enabling Preemptive Multiprogramming on GPUs,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2014
Earlier work this paper cites.
T. Chen, Z. Du, N. Sun, J. Wang, C. Wu, Y. Chen, and O. Temam, “DianNao: A Small-footprint High-throughput Accelerator for Ubiquitous Machine-learning,” in Proceedings of the International Conference on Architectural Support for Programming Languages and Operation Systems (ASPLOS)
2014
Earlier work this paper cites.
Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun, and O. Temam, “DaDianNao: A Machine-Learning Supercomputer,” in Proceedings of the International Symposium on Microarchitecture (MICRO)
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. Le, “Sequence to Sequence Learning with Neural Networks,” in Proceedings of the International Conference on Neural Information Processing Systems (NIPS)
2014
Earlier work this paper cites.
J. Park, Y. Park, and S. Mahlke, “Chimera: Collaborative Preemption for Multitasking on a Shared GPU,” in Proceedings of the International Conference on Architectural Support for Programming Languages and Operation Systems (ASPLOS)
2015
Earlier work this paper cites.
Z. Du, R. Fasthuber, T. Chen, P. Ienne, L. Li, T. Luo, X. Feng, Y. Chen, and O. Temam, “ShiDianNao: Shifting Vision Processing Closer to the Sensor,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2015
Earlier work this paper cites.
D. Liu, T. Chen, S. Liu, J. Zhou, S. Zhou, O. Temam, X. Feng, X. Zhou, and Y. Chen, “PuDianNao: A Polyvalent Machine Learning Accelerator,” in Proceedings of the International Conference on Architectural Support for Programming Languages and Operation Systems (ASPLOS)
2015
Earlier work this paper cites.
Z. Du, D. Rubin, Y. Chen, L. He, T. Chen, L. Zhang, C. Wu, and O. Temam, “Neuromorphic Accelerators: A Comparison Between Neuroscience and Machine-Learning Approaches,” in Proceedings of the International Symposium on Microarchitecture (MICRO)
2015
Earlier work this paper cites.
US 9747546B2
J. Ross, N. Jouppi, A. Phelps, R. Young, T. Norrie, G. Thorson, and D. Luu, “Neural Network Processor.” Patent, 05 2015 · 2015
Earlier work this paper cites.
US 9697463B2
J. Ross and A. Phelps, “Computing Convolutions Using a Neural Network Processor.” Patent, 05 2015 · 2015
Earlier work this paper cites.
US 9805304B2
J. Ross, “Prefetching Weights for Use in a Neural Network Processor.” Patent, 05 2015 · 2015
Earlier work this paper cites.
US 9747548B2
J. Ross and G. Thorson, “Rotating Data for Neural Network Computations.” Patent, 05 2015 · 2015
Earlier work this paper cites.
Y. Kim, W. Yang, and O. Mutlu, “Ramulator: A Fast and Extensible DRAM Simulator,” 2015
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going Deeper with Convolutions,” in Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR)
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” arXiv preprint
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, “Listen, Attend, and Spell,” in arxiv.org
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Y. Chen, J. Emer, and V. Sze, “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2016
Earlier work this paper cites.
J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, “Cnvlutin: Ineffectual-Neuron-Free Deep Convolutional Neural Network Computing,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2016
Earlier work this paper cites.
Z. Lin, L. Nyland, and H. Zhou, “Enabling Efficient Preemption for SIMT Architectures with Lightweight Context Switching,” in Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis (SC)
2016
Cited alongside, same era.
Q. Chen, H. Yang, J. Mars, and L. Tang, “Baymax: Qos awareness and increased utilization for non-preemptive accelerators in warehouse scale computers,” ACM SIGARCH Computer Architecture News
2016
Cited alongside, same era.
B. Reagen, P. Whatmough, R. Adolf, S. Rama, H. Lee, S. Lee, J. Miguel, H. Lobato, G. Wei, and D. Brooks, “Minerva: Enabling Low-Power, High-Accuracy Deep Neural Network Accelerators,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2016
Cited alongside, same era.
P. Chi, S. Li, C. Xu, T. Zhang, J. Zhao, Y. Liu, Y. Wang, and Y. Xie, “A Novel Processing-in-memory Architecture for Neural Network Computation in ReRAM-based Main Memory,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
P. Whatmough, S. Lee, N. Mulholland, P. Hansen, S. Kodali, D. Brooks, and G. Wei, “DNN ENGINE: A 16nm Sub-uJ Deep Neural Network Inference Accelerator for the Embedded Masses,” in Hot Chips: A Symposium on High Performance Chips
2017
Later among the works it cites.
Baidu, “DeepBench: Benchmarking Deep Learning Operations on Different Hardware.” https://github.com/baidu-research/DeepBench , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
NVIDIA, “GeForce GTX Titan Xp (Pascal).” https://www.nvidia.com/en-us/geforce/products/10series/titan-xp , 2017
2017
Later among the works it cites.
Microsoft, “Microsoft Unveils Project Brainwave for Real-time AI,” 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Y. Chen, T. Krishna, J. Emer, and V. Sze, “Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks,” in Proceedings of the International Solid State Circuits Conference (ISSCC)
2016
Cited alongside, same era.
S. Liu, Z. Du, J. Tao, D. Han, T. Luo, Y. Xie, Y. Chen, and T. Chen, “Cambricon: An Instruction Set Architecture for Neural Networks,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2016
Cited alongside, same era.
A. Shafiee, A. Nag, N. Muralimanohar, R. Balasubramonian, J. P. Strachan, M. Hu, R. S. Williams, and V. Srikumar, “ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2016
Cited alongside, same era.
D. Kim, J. Kung, S. Chai, S. Yalamanchili, and S. Mukhopadhyay, “Neurocube: A Programmable Digital Neuromorphic Architecture with High-Density 3D Memory,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2016
Cited alongside, same era.
R. LiKamWa, Y. Hou, M. Polansky, Y. Gao, and L. Zhong, “RedEye: Analog ConvNet Image Sensor Architecture for Continuous Mobile Vision,” in Proceedings of the International Symposium on Computer Architecture (ISCA)
2016
Cited alongside, same era.
D. Mahajan, J. Park, E. Amaro, H. Sharma, A. Yazdan-bakhsh, J. Kim, and H. Esmaeilzadeh, “TABLA: A unified Template-based Framework for Accelerating Statistical Machine Learning,” in Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA)
2016
Cited alongside, same era.
H. Sharma, J. Park, D. Mahajan, E. Amaro, J. Kim, C. Shao, A. Misra, and H. Esmaeilzadeh, “From High-level Deep Neural Models to FPGAs,” in Proceedings of the International Symposium on Microarchitecture (MICRO)
2016
Cited alongside, same era.
M. Rhu, N. Gimelshein, J. Clemons, A. Zulfiqar, and S. W. Keckler, “vDNN: Virtualized Deep Neural Networks for Scalable, Memory-Efficient Neural Network Design,” in Proceedings of the International Symposium on Microarchitecture (MICRO)
2016
Cited alongside, same era.
2017
Later among the works it cites.
S. Poria, E. Cambria, D. Hazarika, N. Mazumder, A. Zadeh, and L. Morency, “Context-Dependent Sentiment Analysis in User-Generated Videos,” in Proceedings of the ACL (Association for Computational Linguistics)
2017
Later among the works it cites.
D. Britz, A. Goldie, M. Luong, and Q. Le, “Massive exploration of neural machine translation architectures,” arXiv preprint
2017
Later among the works it cites.
NVIDIA, “TensorRT Inference Server User Guide,” 2018
2018
Later among the works it cites.
Google, “TensorFlow Serving for Model Deployment in Production,” 2018
2018
Later among the works it cites.
NVIDIA, “NVIDIA Tesla V100,” 2018
2018
Later among the works it cites.
Kubernetes, “Production-Grade Container Orchestration,” 2018
2018
Later among the works it cites.
D. Moss, S. Krishnan, E. Nurvitadhi, P. Ratuszniak, C. Johnson, J. Sim, A. Mishra, D. Marr, S. Subhaschandra, and P. Leong, “A Customizable Matrix Multiplication Framework for the Intel HARPv2 Xeon+FPGA Platform: A Deep Learning Case Study,” in Proceedings of the ACM International Symposium on Field-Programmable Gate Arrays (FPGA)
2018
Later among the works it cites.
Y. Kwon and M. Rhu, “A Case for Memory-Centric HPC System Architecture for Training Deep Neural Networks,” in IEEE Computer Architecture Letters
2018
Later among the works it cites.
Y. Kwon and M. Rhu, “Beyond the Memory Wall: A Case for Memory-Centric HPC System for Deep Learning,” in Proceedings of the International Symposium on Microarchitecture (MICRO)
2018
Later among the works it cites.
2018
Later among the works it cites.
M. Rhu, M. O’Connor, N. Chatterjee, J. Pool, Y. Kwon, and S. W. Keckler, “Compressing DMA Engine: Leveraging Activation Sparsity for Training Deep Neural Networks,” in Proceedings of the International Symposium on High-Performance Computer Architecture (HPCA)
2018
Later among the works it cites.
A. Samajdar, Y. Zhu, P. Whatmough, M. Mattina, and T. Krishna, “SCALE-Sim: Systolic CNN Accelerator Simulator,” in arxiv.org
2018
Later among the works it cites.
Google, “Cloud TPU.” https://cloud.google.com/tpu , 2018
2018
Later among the works it cites.
J. Park, M. Naumov, P. Basu, S. Deng, A. Kalaiah, D. Khudia, J. Law, P. Malani, A. Malevich, S. Nadathur, J. Pino, M. Schatz, A. Sidorov, V. Sivakumar, A. Tulloch, X. Wang, Y. Wu, H. Yuen, U. Diril, D. Dzhulgakov, K. H. an B. Jia, Y. Jia, L. Qiao, V. Rao, N. Rotem, S. Yoo, and M. Smelyanskiy, “Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications,” in arxiv.org
2018
Later among the works it cites.
NVIDIA, “NVIDIA TensorRT: Programmable Inference Accelerator,” 2018
2018
Later among the works it cites.
NVIDIA, “GeForce GTX Titan V (Volta).” https://www.nvidia.com/en-us/titan/titan-v/ , 2018
2018
Later among the works it cites.
NVIDIA, “GeForce GTX 1070.” https://www.nvidia.com/en-in/geforce/products/10series/geforce-gtx-1070/ , 2018
2018
Later among the works it cites.
Intel-Nervana, “Intel Nervana Hardware: Neural Network Processor (Lake Crest),” 2018
2018
Later among the works it cites.
Towards Data Science, “A Beginner’s Guide on Sentiment Analysis with RNN,” 2018
2018
Later among the works it cites.
Harvard NLP, “Open-source Neural Machine Translation.” http://opennmt.net , 2018
2018
Later among the works it cites.
NVIDIA, “NVIDIA AI Inference Platform,” 2018
2018
Later among the works it cites.
T. Singhal, “Maximizing GPU Utilization For Datacenter Inference with NVIDIA TensorRT Inference Server,” 2019
2019
Closest in time.
Google, “Google Cloud: Online versus Batch Prediction.” https://cloud.google.com/ml-engine/docs/tensorflow/online-vs-batch-prediction , 2019
2019
Closest in time.
Y. Kwon, Y. Lee, and M. Rhu, “TensorDIMM: A Practical Near-Memory Processing Architecture for Embeddings and Tensor Operations in Deep Learning,” in Proceedings of the International Symposium on Microarchitecture (MICRO)
2019
Closest in time.
MLPerf, “MLPerf: A broad ML benchmark suite for measuring performance of ML software frameworks, ML hardware accelerators, and ML cloud platforms.” https://github.com/mlperf/inference/tree/master/cloud , 2019
2019
Closest in time.
J. Hestness, N. Ardalani, and G. Diamos, “Beyond Human-Level Accuracy: Computational Challenges in Deep Learning,” in Proceedings of the Symposium on Principles and Practice of Parallel Programming (PPOPP)
2019
Closest in time.