Fetching the paper…
Reading the bibliography…
Large deep neural network (DNN) models pose the key challenge to energy efficiency due to the significantly higher energy consumption of off-chip DRAM accesses than arithmetic or SRAM operations.
1907
Earlier work this paper cites.
S. J. Pan and Q. Yang, “A survey on transfer learning,” IEEE Transactions on knowledge and data engineering , vol. 22, no. 10, pp. 1345–1359, 2010
2010
Earlier work this paper cites.
S. Boyd, N. Parikh, E. Chu, B. Peleato, J. Eckstein et al. , “Distributed optimization and statistical learning via the alternating direction method of multipliers,” Foundations and Trends® in Machine learning , vol. 3, no. 1, pp. 1–122, 2011
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
H. Ouyang, N. He, L. Tran, and A. Gray, “Stochastic alternating direction method of multipliers,” in International Conference on Machine Learning , 2013, pp. 80–88
2013
Earlier work this paper cites.
T. Suzuki, “Dual averaging and proximal gradient descent for online alternating direction multiplier method,” in International Conference on Machine Learning , 2013, pp. 392–400
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
T. Chen, Z. Du, N. Sun, J. Wang, C. Wu, Y. Chen, and O. Temam, “Diannao: A small-footprint high-throughput accelerator for ubiquitous machine-learning,” ACM Sigplan Notices , vol. 49, pp. 269–284, 2014
2014
Earlier work this paper cites.
Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun et al. , “Dadiannao: A machine-learning supercomputer,” in Proceedings of the 47th Annual IEEE/ACM International Symposium on Microarchitecture . IEEE Computer Society, 2014, pp. 609–622
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?” in Advances in neural information processing systems , 2014, pp. 3320–3328
2014
Earlier work this paper cites.
D. Liu, T. Chen, S. Liu, J. Zhou, S. Zhou, O. Teman, X. Feng, X. Zhou, and Y. Chen, “Pudiannao: A polyvalent machine learning accelerator,” in ACM SIGARCH Computer Architecture News , vol. 43, no. 1. ACM, 2015, pp. 369–381
2015
Earlier work this paper cites.
C. Zhang, P. Li, G. Sun, Y. Guan, B. Xiao, and J. Cong, “Optimizing fpga-based accelerator design for deep convolutional neural networks,” in Proceedings of the 2015 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2015, pp. 161–170
2015
Earlier work this paper cites.
Z. Du, R. Fasthuber, T. Chen, P. Ienne, L. Li, T. Luo, X. Feng, Y. Chen, and O. Temam, “Shidiannao: Shifting vision processing closer to the sensor,” in Computer Architecture (ISCA), 2015 ACM/IEEE 42nd Annual International Symposium on . IEEE, 2015, pp. 92–104
2015
Earlier work this paper cites.
S. Han, J. Pool, J. Tran, and W. Dally, “Learning both weights and connections for efficient neural network,” in Advances in neural information processing systems , 2015, pp. 1135–1143
2015
Earlier work this paper cites.
M. Courbariaux, Y. Bengio, and J.-P. David, “Binaryconnect: Training deep neural networks with binary weights during propagations,” in Advances in neural information processing systems , 2015, pp. 3123–3131
2015
Earlier work this paper cites.
H. Sharma, J. Park, D. Mahajan, E. Amaro, J. K. Kim, C. Shao, A. Mishra, and H. Esmaeilzadeh, “From high-level deep neural models to fpgas,” in The 49th Annual IEEE/ACM International Symposium on Microarchitecture . IEEE Press, 2016, p. 17
2016
Earlier work this paper cites.
P. Chi, S. Li, C. Xu, T. Zhang, J. Zhao, Y. Liu, Y. Wang, and Y. Xie, “Prime: A novel processing-in-memory architecture for neural network computation in reram-based main memory,” in ACM SIGARCH Computer Architecture News , vol. 44, no. 3. IEEE Press, 2016, pp. 27–39
2016
Earlier work this paper cites.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “Eie: efficient inference engine on compressed deep neural network,” in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2016, pp. 243–254
2016
Earlier work this paper cites.
J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, “Cnvlutin: Ineffectual-neuron-free deep neural network computing,” ACM SIGARCH Computer Architecture News , vol. 44, no. 3, pp. 1–13, 2016
2016
Earlier work this paper cites.
N. Suda, V. Chandra, G. Dasika, A. Mohanty, Y. Ma, S. Vrudhula, J.-s. Seo, and Y. Cao, “Throughput-optimized opencl-based fpga accelerator for large-scale convolutional neural networks,” in Proceedings of the 2016 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2016, pp. 16–25
2016
Earlier work this paper cites.
J. Qiu, J. Wang, S. Yao, K. Guo, B. Li, E. Zhou, J. Yu, T. Tang, N. Xu, S. Song et al. , “Going deeper with embedded fpga platform for convolutional neural network,” in Proceedings of the 2016 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2016, pp. 26–35
2016
Earlier work this paper cites.
P. Judd, J. Albericio, T. Hetherington, T. M. Aamodt, and A. Moshovos, “Stripes: Bit-serial deep neural network computing,” in Proceedings of the 49th Annual IEEE/ACM International Symposium on Microarchitecture . IEEE Computer Society, 2016, pp. 1–12
2016
Earlier work this paper cites.
B. Reagen, P. Whatmough, R. Adolf, S. Rama, H. Lee, S. K. Lee, J. M. Hernández-Lobato, G.-Y. Wei, and D. Brooks, “Minerva: Enabling low-power, highly-accurate deep neural network accelerators,” in Computer Architecture (ISCA), 2016 ACM/IEEE 43rd Annual International Symposium on . IEEE, 2016, pp. 267–278
2016
Earlier work this paper cites.
D. Mahajan, J. Park, E. Amaro, H. Sharma, A. Yazdanbakhsh, J. K. Kim, and H. Esmaeilzadeh, “Tabla: A unified template-based framework for accelerating statistical machine learning,” in High Performance Computer Architecture (HPCA), 2016 IEEE International Symposium on . IEEE, 2016, pp. 14–26
2016
Earlier work this paper cites.
J. Sim, J.-S. Park, M. Kim, D. Bae, Y. Choi, and L.-S. Kim, “14.6 a 1.42 tops/w deep convolutional neural network recognition processor for intelligent ioe systems,” in Solid-State Circuits Conference (ISSCC), 2016 IEEE International . IEEE, 2016, pp. 264–265
2016
Earlier work this paper cites.
C. Zhang, Z. Fang, P. Zhou, P. Pan, and J. Cong, “Caffeine: towards uniformed representation and acceleration for deep convolutional neural networks,” in Proceedings of the 35th International Conference on Computer-Aided Design . ACM, 2016, p. 12
2016
Earlier work this paper cites.
C. Zhang, D. Wu, J. Sun, G. Sun, G. Luo, and J. Cong, “Energy-efficient cnn implementation on a deeply pipelined fpga cluster,” in Proceedings of the 2016 International Symposium on Low Power Electronics and Design . ACM, 2016, pp. 326–331
2016
Earlier work this paper cites.
https://www.sdxcentral.com/articles/news/intels-deep-learning-chips-will-arrive-2017/2016/11/
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Y. Guo, A. Yao, and Y. Chen, “Dynamic network surgery for efficient dnns,” in Advances In Neural Information Processing Systems , 2016, pp. 1379–1387
2016
Earlier work this paper cites.
D. Lin, S. Talathi, and S. Annapureddy, “Fixed point quantization of deep convolutional networks,” in International Conference on Machine Learning , 2016, pp. 2849–2858
2016
Earlier work this paper cites.
J. Wu, C. Leng, Y. Wang, Q. Hu, and J. Cheng, “Quantized convolutional neural networks for mobile devices,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 4820–4828
2016
Earlier work this paper cites.
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi, “Xnor-net: Imagenet classification using binary convolutional neural networks,” in European Conference on Computer Vision . Springer, 2016, pp. 525–542
2016
Earlier work this paper cites.
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Binarized neural networks,” in Advances in neural information processing systems , 2016, pp. 4107–4115
2016
Earlier work this paper cites.
S. Zhang, Z. Du, L. Zhang, H. Lan, S. Liu, L. Li, Q. Guo, T. Chen, and Y. Chen, “Cambricon-x: An accelerator for sparse neural networks,” in The 49th Annual IEEE/ACM International Symposium on Microarchitecture . IEEE Press, 2016, p. 20
2016
Earlier work this paper cites.
M. Hong, Z.-Q. Luo, and M. Razaviyayn, “Convergence analysis of alternating direction method of multipliers for a family of nonconvex problems,” SIAM Journal on Optimization , vol. 26, no. 1, pp. 337–364, 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
Google supercharges machine learning tasks with TPU custom chip, https://cloudplatform.googleblog.com/2016/05/Google-supercharges-machine-learning-tasks-with-custom-chip.html
2016
Earlier work this paper cites.
S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,” in International Conference on Learning Representations (ICLR) , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Cited alongside, same era.
K. Weiss, T. M. Khoshgoftaar, and D. Wang, “A survey of transfer learning,” Journal of Big Data , vol. 3, no. 1, p. 9, 2016
2016
Cited alongside, same era.
M. Gao, J. Pu, X. Yang, M. Horowitz, and C. Kozyrakis, “Tetris: Scalable and efficient neural network acceleration with 3d memory,” ACM SIGOPS Operating Systems Review , vol. 51, no. 2, pp. 751–764, 2017
2017
Cited alongside, same era.
A. Ren, Z. Li, C. Ding, Q. Qiu, Y. Wang, J. Li, X. Qian, and B. Yuan, “Sc-dcnn: Highly-scalable deep convolutional neural network using stochastic computing,” ACM SIGOPS Operating Systems Review , vol. 51, no. 2, pp. 405–418, 2017
2017
Cited alongside, same era.
C. Gao, D. Neil, E. Ceolini, S.-C. Liu, and T. Delbruck, “Deltarnn: A power-efficient recurrent neural network accelerator,” in Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2018, pp. 21–30
2018
Later among the works it cites.
J. Shen, Y. Huang, Z. Wang, Y. Qiao, M. Wen, and C. Zhang, “Towards a uniform template-based architecture for accelerating 2d and 3d cnns on fpga,” in Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2018, pp. 97–106
2018
Later among the works it cites.
H. Zeng, R. Chen, C. Zhang, and V. Prasanna, “A framework for generating high throughput cnn implementations on fpgas,” in Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2018, pp. 117–126
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Zhao, W. Song, W. Zhang, T. Xing, J.-H. Lin, M. Srivastava, R. Gupta, and Z. Zhang, “Accelerating binarized convolutional neural networks with software-programmable fpgas,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2017, pp. 15–24
2017
Cited alongside, same era.
J. Zhang and J. Li, “Improving the performance of opencl-based fpga accelerator for convolutional neural network,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2017, pp. 25–34
2017
Cited alongside, same era.
C. Zhang and V. Prasanna, “Frequency domain acceleration of convolutional neural networks on cpu-fpga shared memory system,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2017, pp. 35–44
2017
Cited alongside, same era.
Y. Ma, Y. Cao, S. Vrudhula, and J.-s. Seo, “Optimizing loop operation and dataflow in fpga acceleration of deep convolutional neural networks,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2017, pp. 45–54
2017
Cited alongside, same era.
U. Aydonat, S. O’Connell, D. Capalija, A. C. Ling, and G. R. Chiu, “An opencl™ deep learning accelerator on arria 10,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2017, pp. 55–64
2017
Cited alongside, same era.
Y. Umuroglu, N. J. Fraser, G. Gambardella, M. Blott, P. Leong, M. Jahre, and K. Vissers, “Finn: A framework for fast, scalable binarized neural network inference,” in Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2017, pp. 65–74
2017
Cited alongside, same era.
S. Venkataramani, A. Ranjan, S. Banerjee, D. Das, S. Avancha, A. Jagannathan, A. Durg, D. Nagaraj, B. Kaul, P. Dubey et al. , “Scaledeep: A scalable compute architecture for learning and evaluating deep networks,” in Computer Architecture (ISCA), 2017 ACM/IEEE 44th Annual International Symposium on . IEEE, 2017, pp. 13–26
2017
Cited alongside, same era.
Y.-H. Chen, T. Krishna, J. S. Emer, and V. Sze, “Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks,” IEEE Journal of Solid-State Circuits , vol. 52, no. 1, pp. 127–138, 2017
2017
Cited alongside, same era.
2018
Later among the works it cites.
Z. Chen, A. Howe, H. T. Blair, and J. Cong, “Fpga-based lstm acceleration for real-time eeg signal processing,” in Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2018, pp. 288–288
2018
Later among the works it cites.
Y. Du, Q. Liu, S. Wei, and C. Gao, “Software-defined fpga-based accelerator for deep convolutional neural networks,” in Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2018, pp. 291–291
2018
Later among the works it cites.
S. Liu, X. Niu, and W. Luk, “A low-power deconvolutional accelerator for convolutional neural network based segmentation on fpga,” in Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2018, pp. 293–293
2018
Later among the works it cites.
M. Song, K. Zhong, J. Zhang, Y. Hu, D. Liu, W. Zhang, J. Wang, and T. Li, “In-situ ai: Towards autonomous and incremental deep learning for iot systems,” in High Performance Computer Architecture (HPCA), 2018 IEEE International Symposium on . IEEE, 2018, pp. 92–103
2018
Later among the works it cites.
2018
Later among the works it cites.
Z. Yuan, J. Yue, H. Yang, Z. Wang, J. Li, Y. Yang, Q. Guo, X. Li, M.-F. Chang, H. Yang et al. , “Sticker: A 0.41-62.1 tops/w 8bit neural network processor with multi-sparsity compatible convolution arrays and online tuning acceleration for fully connected layers,” in 2018 IEEE Symposium on VLSI Circuits . IEEE, 2018, pp. 33–34
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
S. Liu, J. Chen, P.-Y. Chen, and A. Hero, “Zeroth-order online alternating direction method of multipliers: Convergence analysis and applications,” in International Conference on Artificial Intelligence and Statistics , 2018, pp. 288–297
2018
Later among the works it cites.
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “Mobilenetv2: Inverted residuals and linear bottlenecks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 4510–4520
2018
Later among the works it cites.
H.-P. Cheng, Y. Huang, X. Guo, F. Yan, Y. Huang, W. Wen, H. Li, and Y. Chen, “Differentiable fine-grained quantization for deep neural network compression,” in NIPS 2018 CDNNRIA Workshop , 2018
2018
Later among the works it cites.
R. Yu, A. Li, C.-F. Chen, J.-H. Lai, V. I. Morariu, X. Han, M. Gao, C.-Y. Lin, and L. S. Davis, “Nisp: Pruning networks using neuron importance score propagation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 9194–9203
2018
Later among the works it cites.
Z. Zhuang, M. Tan, B. Zhuang, J. Liu, Y. Guo, Q. Wu, J. Huang, and J. Zhu, “Discrimination-aware channel pruning for deep neural networks,” in Advances in Neural Information Processing Systems , 2018, pp. 875–886
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
Y. He, J. Lin, Z. Liu, H. Wang, L.-J. Li, and S. Han, “Amc: Automl for model compression and acceleration on mobile devices,” in The European Conference on Computer Vision (ECCV) , September 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2019
Closest in time.
X. Wang, J. Yu, C. Augustine, R. Iyer, and R. Das, “Bit prudent in-cache acceleration of deep convolutional neural networks,” in 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2019, pp. 81–93
2019
Closest in time.
Y. Yang, Q. Huang, B. Wu, T. Zhang, L. Ma, G. Gambardella, M. Blott, L. Lavagno, K. Vissers, J. Wawrzynek et al. , “Synetgy: Algorithm-hardware co-design for convnet accelerators on embedded fpgas,” in Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2019, pp. 23–32
2019
Closest in time.
L. Jing, J. Liu, and F. Yu, “A deep learning inference accelerator based on model compression on fpga,” in Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2019, pp. 118–118
2019
Closest in time.
W. You and C. Wu, “A reconfigurable accelerator for sparse convolutional neural networks,” in Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2019, pp. 119–119
2019
Closest in time.
X. Wei, Y. Liang, P. Zhang, C. H. Yu, and J. Cong, “Overcoming data transfer bottlenecks in dnn accelerators via layer-conscious memory managment,” in Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2019, pp. 120–120
2019
Closest in time.
J. Zhang and J. Li, “Unleashing the power of soft logic for convolutional neural network acceleration via product quantization,” in Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2019, pp. 120–120
2019
Closest in time.
S. Zeng, Y. Lin, S. Liang, J. Kang, D. Xie, Y. Shan, S. Han, Y. Wang, and H. Yang, “A fine-grained sparse accelerator for multi-precision dnn,” in Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2019, pp. 185–185
2019
Closest in time.
H. Nakahara, A. Jinguji, M. Shimoda, and S. Sato, “An fpga-based fine tuning accelerator for a sparse cnn,” in Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2019, pp. 186–186
2019
Closest in time.
L. Lu, Y. Liang, R. Huang, W. Lin, X. Cui, and J. Zhang, “Speedy: An accelerator for sparse convolutional neural networks on fpgas,” in Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2019, pp. 187–187
2019
Closest in time.
Z. Tang, G. Luo, and M. Jiang, “Ftconv: Fpga acceleration for transposed convolution layers in deep neural networks,” in Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2019, pp. 189–189
2019
Closest in time.
K. Guo, S. Liang, J. Yu, X. Ning, W. Li, Y. Wang, and H. Yang, “Compressed cnn training with fpga-based accelerator,” in Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2019, pp. 189–189
2019
Closest in time.
E. Wu, X. Zhang, D. Berman, I. Cho, and J. Thendean, “Compute-efficient neural-network acceleration,” in Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2019, pp. 191–200
2019
Closest in time.
S. Vogel, J. Springer, A. Guntoro, and G. Ascheid, “Efficient acceleration of cnns for semantic segmentation on fpgas,” in Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays . ACM, 2019, pp. 309–309
2019
Closest in time.
A. Ren, J. Li, T. Zhang, S. Ye, W. Xu, X. Qian, X. Lin, and Y. Wang, “ADMM-NN: An Algorithm-Hardware Co-Design Framework of DNNs Using Alternating Direction Methods of Multipliers,” in International conference on Architectural Support for Programming Languages and Operating Systems , 2019
2019
Closest in time.
A. Ren, T. Zhang, S. Ye, J. Li, W. Xu, X. Qian, X. Lin, and Y. Wang, “Admm-nn: An algorithm-hardware co-design framework of dnns using alternating direction methods of multipliers,” in Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems . ACM, 2019, pp. 925–938
2019
Closest in time.
2019
Closest in time.
W. Wen, C. Wu, Y. Wang, Y. Chen, and H. Li, “Learning structured sparsity in deep neural networks,” in Advances in Neural Information Processing Systems , 2016, pp. 2074–2082
2082
Closest in time.