Fetching the paper…
Reading the bibliography…
Convolutional neural network (CNN) inference on mobile devices demands efficient hardware acceleration of low-precision (INT8) general matrix multiplication (GEMM).
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database,” in CVPR09 , 2009
2009
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” Master’s thesis, Department of Computer Science, University of Toronto , 2009
2009
Earlier work this paper cites.
Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun, and O. Temam, “Dadiannao: A machine-learning supercomputer,” in Proceedings of the 47th Annual IEEE/ACM International Symposium on Microarchitecture , ser. MICRO-47. Washington, DC, USA: IEEE Computer Society, 2014, pp. 609–622. [Online]. Available: http://dx.doi.org.ezp-prod1.hul.harvard.edu/10.1109/MICRO.2014.58
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
P. Warden. Why GEMM is at the heard of deep learning. [Online]. Available: https://petewarden.com/2015/04/20/why-gemm-is-at-the-heart-of-deep-learning/
2015
Earlier work this paper cites.
J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, “Cnvlutin: Ineffectual-neuron-free deep neural network computing,” in Proceedings of the 43rd International Symposium on Computer Architecture , ser. ISCA ’16. Piscataway, NJ, USA: IEEE Press, 2016, pp. 1–13. [Online]. Available: https://doi.org/10.1109/ISCA.2016.11
2016
Earlier work this paper cites.
Y.-H. Chen, J. Emer, and V. Sze, “Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks,” in Proceedings of the 43rd International Symposium on Computer Architecture . Piscataway, NJ, USA: IEEE Press, 2016, pp. 367–379. [Online]. Available: https://doi.org/10.1109/ISCA.2016.40
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Han et al. , “EIE: Efficient inference engine on compressed deep neural network,” in Proceedings of the 43rd Int. Symp. on Computer Architecture (ISCA) , 2016
2016
Earlier work this paper cites.
B. Reagen, P. Whatmough, R. Adolf, S. Rama, H. Lee, S. K. Lee, J. M. Hernández-Lobato, G.-Y. Wei, and D. Brooks, “Minerva: Enabling low-power, highly-accurate deep neural network accelerators,” in Proceedings of the 43rd International Symposium on Computer Architecture , ser. ISCA, 2016
2016
Earlier work this paper cites.
J. Albericio, A. Delmás, P. Judd, S. Sharify, G. O’Leary, R. Genov, and A. Moshovos, “Bit-pragmatic deep neural network computing,” in Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture , ser. MICRO-50 ’17. New York, NY, USA: Association for Computing Machinery, 2017, p. 382–394. [Online]. Available: https://doi.org/10.1145/3123939.3123982
2017
Earlier work this paper cites.
S. Han, J. Kang, H. Mao, Y. Hu, X. Li, Y. Li, D. Xie, H. Luo, S. Yao, Y. Wang, H. Yang, and W. B. J. Dally, “Ese: Efficient speech recognition engine with sparse lstm on fpga,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays , ser. FPGA ’17. New York, NY, USA: Association for Computing Machinery, 2017, p. 75–84. [Online]. Available: https://doi.org/10.1145/3020078.3021745
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
S. Kodali, P. Hansen, N. Mulholland, P. Whatmough, D. Brooks, and G. Wei, “Applications of Deep Neural Networks for Ultra Low Power IoT,” in 2017 IEEE International Conference on Computer Design (ICCD) , 2017, pp. 589–592
2017
Cited alongside, same era.
A. Parashar et al. , “SCNN: An accelerator for compressed-sparse convolutional neural networks,” in 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA) , June 2017, pp. 27–40
2017
Cited alongside, same era.
H. Kang, “Accelerator-aware pruning for convolutional neural networks,” IEEE Transactions on Circuits and Systems for Video Technology , pp. 1–1, 2019
2019
Later among the works it cites.
H. Kung, B. McDanel, and S. Q. Zhang, “Packing sparse convolutional neural networks for efficient systolic array implementations: Column combining under joint optimization,” in 24th Int. Conf. on Architectural Support for Programming Languages and Operating Systems (ASPLOS) , 2019, pp. 821–834
2019
Later among the works it cites.
H. Li, M. Bhargav, P. N. Whatmough, and H. . Philip Wong, “On-chip memory technology design space explorations for mobile deep neural network accelerators,” in 2019 56th ACM/IEEE Design Automation Conference (DAC) , 2019, pp. 1–6
2019
Later among the works it cites.
S. Sharify, A. D. Lascorz, M. Mahmoud, M. Nikolic, K. Siu, D. M. Stuart, Z. Poulos, and A. Moshovos, “Laconic deep learning inference acceleration,” in Proceedings of the 46th International Symposium on Computer Architecture , ser. ISCA ’19. New York, NY, USA: Association for Computing Machinery, 2019, p. 304–317. [Online]. Available: https://doi.org/10.1145/3307650.3322255
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
C. Deng, S. Liao, Y. Xie, K. K. Parhi, X. Qian, and B. Yuan, “Permdnn: Efficient compressed dnn architecture with permuted diagonal matrices,” in 2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2018, pp. 189–202
2018
Cited alongside, same era.
S. Pal et al. , “Outerspace: An outer product based sparse matrix multiplication accelerator,” in Int. Symp. on High Performance Computer Architecture (HPCA) , Feb 2018, pp. 724–736
2018
Cited alongside, same era.
H. Sharma, J. Park, N. Suda, L. Lai, B. Chau, V. Chandra, and H. Esmaeilzadeh, “Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural networks,” in Proceedings of the 45th Annual International Symposium on Computer Architecture , ser. ISCA ’18. IEEE Press, 2018, p. 764–775. [Online]. Available: https://doi.org/10.1109/ISCA.2018.00069
2018
Cited alongside, same era.
X. Zhou, Z. Du, Q. Guo, S. Liu, C. Liu, C. Wang, X. Zhou, L. Li, T. Chen, and Y. Chen, “Cambricon-s: Addressing irregularity in sparse neural networks through a cooperative software/hardware approach,” in Proceedings of the 51st Annual IEEE/ACM International Symposium on Microarchitecture , ser. MICRO-51. IEEE Press, 2018, p. 15–28. [Online]. Available: https://doi.org/10.1109/MICRO.2018.00011
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Y. Chen, T. Yang, J. Emer, and V. Sze, “Eyeriss v2: A flexible accelerator for emerging deep neural networks on mobile devices,” IEEE Journal on Emerging and Selected Topics in Circuits and Systems , vol. 9, no. 2, pp. 292–308, 2019
2019
Cited alongside, same era.
I. Fedorov, R. P. Adams, M. Mattina, and P. N. Whatmough, “SpArSe: Sparse architecture search for CNNs on resource-constrained microcontrollers,” in Advances in Neural Information Processing Systems (NeurIPS) , 2019, pp. 4978–4990
2019
Cited alongside, same era.
Y. Feng, P. Whatmough, and Y. Zhu, “ASV: Accelerated Stereo Vision System,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , ser. MICRO ’52. New York, NY, USA: Association for Computing Machinery, 2019, p. 643–656. [Online]. Available: https://doi.org/10.1145/3352460.3358253
2019
Cited alongside, same era.
2019
Later among the works it cites.
G. Shomron, T. Horowitz, and U. Weiser, “SMT-SA: Simultaneous multithreading in systolic arrays,” IEEE Comput. Archit. Lett. , vol. 18, no. 2, pp. 99–102, Jul. 2019
2019
Later among the works it cites.
P. N. Whatmough, S. K. Lee, M. Donato, H. Hsueh, S. Xi, U. Gupta, L. Pentecost, G. G. Ko, D. Brooks, and G. Wei, “A 16nm 25mm2 SoC with a 54.5x Flexibility-Efficiency Range from Dual-Core Arm Cortex-A53 to eFPGA and Cache-Coherent Accelerators,” in 2019 Symposium on VLSI Circuits , 2019, pp. C34–C35
2019
Later among the works it cites.
P. N. Whatmough, C. Zhou, P. Hansen, S. K. Venkataramanaiah, J. sun Seo, and M. Mattina, “FixyNN: Efficient Hardware for Mobile Computer Vision via Transfer Learning,” in Proceedings of the 2nd SysML Conference, Palo Alto, CA, USA , 2019
2019
Later among the works it cites.
I. Fedorov, M. Stamenovic, C. Jenson, L.-C. Yang, A. Mandell, Y. Gan, M. Mattina, and P. N. Whatmough, “ TinyLSTMs: Efficient Neural Speech Enhancement for Hearing Aids ,” in Conference of the International Speech Communication Association (INTERSPEECH) , 2020
2020
Closest in time.
P. Hansen, A. Vilkin, Y. Khrustalev, J. Imber, D. Hanwell, M. Mattina, and P. N. Whatmough, “ ISP4ML: Understanding the Role of Image Signal Processing in Efficient Deep Learning Vision Systems ,” in International Conference on Pattern Recognition (ICPR) , 2020
2020
Closest in time.
Z. Liu, P. N. Whatmough, and M. Mattina, “Systolic Tensor Array: An Efficient Structured-Sparse GEMM Accelerator for Mobile CNN Inference,” IEEE Computer Architecture Letters , vol. 19, no. 1, pp. 34–37, 2020
2020
Closest in time.
X. Yang, M. Gao, Q. Liu, J. Setter, J. Pu, A. Nayak, S. Bell, K. Cao, H. Ha, P. Raina, C. Kozyrakis, and M. Horowitz, “Interstellar: Using halide’s scheduling language to analyze dnn accelerators,” in Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems , ser. ASPLOS ’20. New York, NY, USA: Association for Computing Machinery, 2020, p. 369–383. [Online]. Available: https://doi.org/10.1145/3373376.3378514
2020
Closest in time.
Z. Zhang, H. Wang, S. Han, and W. J. Dally, “Sparch: Efficient architecture for sparse matrix multiplication,” in 26th IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2020
2020
Closest in time.