Fetching the paper…
Reading the bibliography…
Network pruning can reduce the high computation cost of deep neural network (DNN) models.
1906
Earlier work this paper cites.
Y. LeCun, J. S. Denker, and S. A. Solla, “Optimal brain damage,” in Advances in neural information processing systems , 1990, pp. 598–605
1990
Earlier work this paper cites.
Y. LeCun, Y. Bengio et al. , “Convolutional networks for images, speech, and time series,” The handbook of brain theory and neural networks , vol. 3361, no. 10, p. 1995, 1995
1995
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting on association for computational linguistics . Association for Computational Linguistics, 2002, pp. 311–318
2002
Earlier work this paper cites.
K. Chellapilla, S. Puri, and P. Simard, “High performance convolutional neural networks for document processing,” 2006
2006
Earlier work this paper cites.
J. Leng, T. Hetherington, A. ElTantawy, S. Gilani, N. S. Kim, T. M. Aamodt, and V. J. Reddi, “GPUWattch: Enabling Energy Optimizations in GPGPUs,” in Proceedings of the 40th Annual International Symposium on Computer Architecture . New York, NY, USA: Association for Computing Machinery, 2013, pp. 487–498. [Online]. Available: https://doi.org/10.1145/2485922.2485964
2013
Earlier work this paper cites.
M. Rhu, M. Sullivan, J. Leng, and M. Erez, “A locality-aware memory hierarchy for energy-efficient gpu architectures,” in 2013 46th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2013, pp. 86–98
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, software available from tensorflow.org. [Online]. Available: https://www.tensorflow.org/
2015
Earlier work this paper cites.
S. Han, J. Pool, J. Tran, and W. Dally, “Learning both weights and connections for efficient neural network,” in Advances in neural information processing systems , 2015, pp. 1135–1143
2015
Earlier work this paper cites.
NVIDIA, “GPU Pro Tip: CUDA 7 Streams Simplify Concurrency,” https://devblogs.nvidia.com/gpu-pro-tip-cuda-7-streams-simplify-concurrency/, 2015
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in ICLR 2015 : International Conference on Learning Representations 2015 , 2015
2015
Earlier work this paper cites.
Y.-H. Chen, T. Krishna, J. S. Emer, and V. Sze, “Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks,” IEEE journal of solid-state circuits , vol. 52, no. 1, pp. 127–138, 2016
2016
Earlier work this paper cites.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “Eie: Efficient inference engine on compressed deep neural network,” in Proceedings of the 43rd International Symposium on Computer Architecture , ser. ISCA ’16. IEEE Press, 2016, p. 243–254. [Online]. Available: https://doi.org/10.1109/ISCA.2016.30
2016
Earlier work this paper cites.
M.-T. Luong and C. D. Manning, “Achieving open vocabulary neural machine translation with hybrid word-character models,” in Association for Computational Linguistics (ACL) , Berlin, Germany, August 2016. [Online]. Available: https://nlp.stanford.edu/pubs/luong2016acl_hybrid.pdf
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
P. Hill, A. Jain, M. Hill, B. Zamirai, C.-H. Hsu, M. A. Laurenzano, S. Mahlke, L. Tang, and J. Mars, “Deftnn: Addressing bottlenecks for dnn execution on gpus via synapse vector elimination and near-compute data fission,” in Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture . ACM, 2017, pp. 786–799
2017
Earlier work this paper cites.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers et al. , “In-datacenter performance analysis of a tensor processing unit,” in 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2017, pp. 1–12
2017
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Communications of The ACM , vol. 60, no. 6, pp. 84–90, 2017
2017
Earlier work this paper cites.
J.-H. Luo, J. Wu, and W. Lin, “Thinet: A filter level pruning method for deep neural network compression,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 5058–5066
2017
Earlier work this paper cites.
M. Luong, E. Brevdo, and R. Zhao, “Neural machine translation (seq2seq) tutorial,” https://github.com/tensorflow/nmt , 2017
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
——, “NVIDIA Volta GPU Architecture Whitepaper,” 2017
2017
Cited alongside, same era.
A. Parashar, M. Rhu, A. Mukkara, A. Puglielli, R. Venkatesan, B. Khailany, J. Emer, S. W. Keckler, and W. J. Dally, “Scnn: An accelerator for compressed-sparse convolutional neural networks,” in Proceedings of the 44th Annual International Symposium on Computer Architecture , ser. ISCA ’17. New York, NY, USA: Association for Computing Machinery, 2017, p. 27–40. [Online]. Available: https://doi.org/10.1145/3079856.3080254
2017
Cited alongside, same era.
2017
Cited alongside, same era.
P. Molchanov, A. Mallya, S. Tyree, I. Frosio, and J. Kautz, “Importance estimation for neural network pruning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 11 264–11 272
2019
Later among the works it cites.
I. Nisa, J. Li, A. Sukumaran-Rajam, P. S. Rawat, S. Krishnamoorthy, and P. Sadayappan, “An efficient mixed-mode representation of sparse tensors,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2019, pp. 1–25
2019
Later among the works it cites.
——, “CUDA Toolkit Documentation v10.1,” 2019
2019
Later among the works it cites.
——, “CUTLASS 1.3,” https://github.com/NVIDIA/cutlass, 2019
2019
Later among the works it cites.
Y. Qiu, J. Leng, C. Guo, Q. Chen, C. Li, M. Guo, and Y. Zhu, “Adversarial Defense Through Network Profiling Based Path Extraction,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 4777–4786
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Cited alongside, same era.
T.-J. Yang, Y.-H. Chen, and V. Sze, “Designing energy-efficient convolutional neural networks using energy-aware pruning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 5687–5695
2017
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
U. A. ROMAN STEINBERG, “6 areas where artificial neural networks outperform humans,” https://venturebeat.com/2017/12/08/6-areas-where-artificial-neural-networks-outperform-humans/, 2018
2018
Cited alongside, same era.
T.-J. Yang, A. Howard, B. Chen, X. Zhang, A. Go, M. Sandler, V. Sze, and H. Adam, “Netadapt: Platform-aware neural network adaptation for mobile applications,” in Proceedings of the European Conference on Computer Vision , 2018, pp. 285–300
2018
Cited alongside, same era.
R. Yu, A. Li, C.-F. Chen, J.-H. Lai, V. I. Morariu, X. Han, M. Gao, C.-Y. Lin, and L. S. Davis, “Nisp: Pruning networks using neuron importance score propagation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 9194–9203
2018
Cited alongside, same era.
T. Zhang, S. Ye, K. Zhang, X. Ma, N. Liu, L. Zhang, J. Tang, K. Ma, X. L. Lin, M. Fardad, and Y. Wang, “Structadmm: A systematic, high-efficiency framework of structured weight pruning for dnns,” 2018
2018
Cited alongside, same era.
2019
Later among the works it cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI Blog , vol. 1, no. 8, p. 9, 2019
2019
Later among the works it cites.
M. A. Raihan, N. Goli, and T. M. Aamodt, “Modeling deep learning accelerator enabled gpus,” in 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) , 2019, pp. 79–92
2019
Later among the works it cites.
F. Sadi, J. Sweeney, T. M. Low, J. C. Hoe, L. Pileggi, and F. Franchetti, “Efficient spmv operation for large and highly sparse matrices using scalable multi-way merge parallelization,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , 2019, pp. 347–358
2019
Later among the works it cites.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, “Glue: A multi-task benchmark and analysis platform for natural language understanding,” in ICLR 2019 : 7th International Conference on Learning Representations , 2019
2019
Later among the works it cites.
H. Yang, Y. Zhu, and J. Liu, “Ecc: Energy-constrained deep neural network compression via a bilinear regression model,” International Conference on Computer Vision and Pattern Recognition (CVPR) , 2019
2019
Later among the works it cites.
——, “Energy-constrained compression for deep neural networks via weighted sparse projection and layer input masking,” International Conference on Learning Representations (ICLR) , 2019
2019
Later among the works it cites.
T.-H. Yang, H.-Y. Cheng, C.-L. Yang, I.-C. Tseng, H.-W. Hu, H.-S. Chang, and H.-P. Li, “Sparse reram engine: joint exploration of activation and weight sparsity in compressed neural networks,” in Proceedings of the 46th International Symposium on Computer Architecture , 2019, pp. 236–249
2019
Later among the works it cites.
Z. Yao, S. Cao, W. Xiao, C. Zhang, and L. Nie, “Balanced sparsity for efficient dnn inference on gpu,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, 2019, pp. 5676–5683
2019
Later among the works it cites.
J. Zhang, X. Chen, M. Song, and T. Li, “Eager pruning: algorithm and architecture support for fast training of deep neural networks,” in Proceedings of the 46th International Symposium on Computer Architecture . ACM, 2019, pp. 292–303
2019
Later among the works it cites.
M. Zhu, T. Zhang, Z. Gu, and Y. Xie, “Sparse tensor core: Algorithm and hardware co-design for vector-wise sparse neural networks on modern gpus,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture . ACM, 2019, pp. 359–371
2019
Later among the works it cites.
2020
Closest in time.
C. Guo, Y. Zhou, J. Leng, Y. Zhu, Z. Du, Q. Chen, C. Li, B. Yao, and M. Guo, “Balancing Efficiency and Flexibility for DNN Acceleration via Temporal GPU-Systolic Array Integration,” in 2020 57th ACM/IEEE Design Automation Conference (DAC) , 2020, pp. 1–6
2020
Closest in time.
2020
Closest in time.
W. Niu, X. Ma, S. Lin, S. Wang, X. Qian, X. Lin, Y. Wang, and B. Ren, “Patdnn: Achieving real-time dnn execution on mobile devices with pattern-based weight pruning,” in Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems , ser. ASPLOS ’20, 2020
2020
Closest in time.
E. Qin, A. Samajdar, H. Kwon, V. Nadella, S. Srinivasan, D. Das, B. Kaul, and T. Krishna, “Sigma: A sparse and irregular gemm accelerator with flexible interconnects for dnn training,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2020, pp. 58–70
2020
Closest in time.
P. Tillet, “Torch-Blocksparse,” https://github.com/ptillet/torch-blocksparse, 2020
2020
Closest in time.
H. Yang, S. Gui, Y. Zhu, and J. Liu, “Automatic neural network compression by sparsity-quantization joint learning: A constrained optimization-based approach,” International Conference on Computer Vision and Pattern Recognition (CVPR) , 2020
2020
Closest in time.
W. Wen, C. Wu, Y. Wang, Y. Chen, and H. Li, “Learning structured sparsity in deep neural networks,” in Advances in neural information processing systems , 2016, pp. 2074–2082
2082
Closest in time.