Fetching the paper…
Reading the bibliography…
A growing number of applications implement predictive functions using deep learning models, which require heavy use of compute and memory.
1903
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
2014
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” CoRR , 2015
2015
Earlier work this paper cites.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015
2015
Earlier work this paper cites.
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang, “Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems,” CoRR , 2015
2015
Earlier work this paper cites.
D. Lin, S. Talathi, and S. Annapureddy, “Fixed point quantization of deep convolutional networks,” in International Conference on Machine Learning , 2016, pp. 2849–2858
2016
Earlier work this paper cites.
J. Wu, C. Leng, Y. Wang, Q. Hu, and J. Cheng, “Quantized convolutional neural networks for mobile devices,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 4820–4828
2016
Earlier work this paper cites.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers et al. , “In-datacenter performance analysis of a tensor processing unit,” in Proceedings of the 44th Annual International Symposium on Computer Architecture , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Nvidia, “8 bit inference with TensorRT.” [Online]. Available: http://on-demand.gputechconf.com/gtc/2017/presentation/s7310-8-bit-inferencewith-tensorrt.pdf
2017
Earlier work this paper cites.
2018
Cited alongside, same era.
K. Hazelwood, S. Bird, D. Brooks, S. Chintala, U. Diril, D. Dzhulgakov, M. Fawzy, B. Jia, Y. Jia, A. Kalro et al. , “Applied machine learning at facebook: A datacenter infrastructure perspective,” in 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2018
2018
Cited alongside, same era.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018
2018
Cited alongside, same era.
Z. Jiang, T. Chen, and M. Li, “Efficient deep learning inference on edge devices,” in Proceedings of ACM Conference on Systems and Machine Learning (SysML’18) , 2018
2018
Cited alongside, same era.
L. Wang, Z. Chen, Y. Liu, Y. Wang, L. Zheng, M. Li, and Y. Wang, “A unified optimization approach for cnn model inference on integrated gpus,” in Proceedings of the 48th International Conference on Parallel Processing , 2019, pp. 1–10
2019
Later among the works it cites.
Y. Liu, Y. Wang, R. Yu, M. Li, V. Sharma, and Y. Wang, “Optimizing { \{ CNN } \} model inference on cpus,” in 2019 { \{ USENIX } \} Annual Technical Conference ( { \{ USENIX } \} { \{ ATC } \} 19) , 2019, pp. 1025–1040
2019
Later among the works it cites.
2019
Later among the works it cites.
“Low precision inference on gpus.” [Online]. Available: https://developer.download.nvidia.com/video/ gputechconf/gtc/2019/presentation/s9659-inference-at-reduced-precision-on-gpus.pdf
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, H. Shen, M. Cowan, L. Wang, Y. Hu, L. Ceze, C. Guestrin, and A. Krishnamurthy, “TVM: An automated end-to-end optimizing compiler for deep learning,” in USENIX Symposium on Operating Systems Design and Implementation (OSDI) , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, “Quantization and training of neural networks for efficient integer-arithmetic-only inference,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-performance deep learning library,” in Advances in Neural Information Processing Systems 32 , 2019
2019
Cited alongside, same era.
“TFLite hosted pre-quantized models.” [Online]. Available: https://www.tensorflow.org/lite/guide/hosted_models
Cited in the paper.
C. Lattner and J. Pienaar, “Mlir primer: A compiler infrastructure for the end of moore’s law,” 2019
2019
Later among the works it cites.
“Intel MKLDNN,” https://github.com/intel/mkl-dnn, accessed: 2020-02-10
2020
Closest in time.
“Intel VNNI instruction,” https://software.intel.com/en-us/ai/deep-learning-boost, accessed: 2020-02-10
2020
Closest in time.
“ARM DOTI instruction,” https://community.arm.com/developer/tools-software/tools/b/tools-software-ides-blog/posts/exploring-the-arm-dot-product-instructions, accessed: 2020-02-10
2020
Closest in time.
“XLA: Optimizing compiler for machine learning,” https://www.tensorflow.org/xla, accessed: 2020-02-10
2020
Closest in time.
M. Cowan, T. Moreau, T. Chen, J. Bornholt, and L. Ceze, “Automatic generation of high-performance quantized machine learning kernels,” in Proceedings of the 18th ACM/IEEE International Symposium on Code Generation and Optimization , 2020, pp. 305–316
2020
Closest in time.