Fetching the paper…
Reading the bibliography…
In recent years, there has been a flurry of research in deep neural network pruning and compression.
The sparse polyhedral framework: Composing compiler-generated inspector-executor code
Michelle Mills Strout, Mary Hall, and Catherine Olschanowsky. 2018 · 1934
Earlier work this paper cites.
Automated empirical optimizations of software and the ATLAS project
R Clint Whaley, Antoine Petitet, and Jack J Dongarra. 2001 · 2001
Earlier work this paper cites.
Sparse tiling for stationary iterative methods
Michelle Mills Strout, Larry Carter, Jeanne Ferrante, and Barbara Kreaseck. 2004 · 2004
Earlier work this paper cites.
Opentuner: An extensible framework for program autotuning. In Proceedings of the 23rd international conference on Parallel architectures and compilation . ACM, 303–316
Jason Ansel, Shoaib Kamil, Kalyan Veeramachaneni, Jonathan Ragan-Kelley, Jeffrey Bosboom, Una-May O’Reilly, and Saman Amarasinghe. 2014 · 2014
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014 · 2014
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Loop and data transformations for sparse matrix code. In ACM SIGPLAN Notices , Vol. 50. ACM, 521–532
Anand Venkat, Mary Hall, and Michelle Strout. 2015 · 2015
Earlier work this paper cites.
EIE: efficient inference engine on compressed deep neural network. In 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) . IEEE, 243–254
Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A Horowitz, and William J Dally. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Fast algorithms for convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 4013–4021
Andrew Lavin and Scott Gray. 2016 · 2016
Earlier work this paper cites.
Deepbench
Sharan Narang. 2016 · 2016
Earlier work this paper cites.
Faster cnns with direct sparse convolutions and guided pruning
Jongsoo Park, Sheng Li, Wei Wen, Ping Tak Peter Tang, Hai Li, Yiran Chen, and Pradeep Dubey. 2016 · 2016
Earlier work this paper cites.
Gpu kernels for block-sparse weights
Scott Gray, Alec Radford, and Diederik P Kingma. 2017 · 2017
Earlier work this paper cites.
Cooperative groups: Flexible CUDA thread programming
M Harris and K Perelygin. 2017 · 2017
Earlier work this paper cites.
Channel pruning for accelerating very deep neural networks. In Proceedings of the IEEE International Conference on Computer Vision . 1389–1397
Yihui He, Xiangyu Zhang, and Jian Sun. 2017 · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017 · 2017
Cited alongside, same era.
Learning Sparse Neural Networks through L _ 0 L\_0 Regularization
Christos Louizos, Max Welling, and Diederik P Kingma. 2017 · 2017
Cited alongside, same era.
Variational dropout sparsifies deep neural networks. In Proceedings of the 34th International Conference on Machine Learning-Volume 70 . JMLR. org, 2498–2507
Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov. 2017 · 2017
Cited alongside, same era.
Acorns: A framework for accelerating deep neural networks with input sparsity. In 2019 28th International Conference on Parallel Architectures and Compilation Techniques (PACT) . IEEE, 178–191
Xiao Dong, Lei Liu, Peng Zhao, Guangli Li, Jiansong Li, Xueying Wang, and Xiaobing Feng. 2019 · 2019
Later among the works it cites.
Erich Elsen, Marat Dukhan, Trevor Gale, and Karen Simonyan. 2019 · 2019
Later among the works it cites.
Rigging the Lottery: Making All Tickets Winners
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen. 2019 · 2019
Later among the works it cites.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Input-aware auto-tuning of compute-bound HPC kernels. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . ACM, 43
Philippe Tillet and David Cox. 2017 · 2017
Cited alongside, same era.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Learning intrinsic sparse structures within long short-term memory
Wei Wen, Yuxiong He, Samyam Rajbhandari, Minjia Zhang, Wenhan Wang, Fang Liu, Bin Hu, Yiran Chen, and Hai Li. 2017 · 2017
Cited alongside, same era.
TVM: end-to-end optimization stack for deep learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Haichen Shen, Eddie Q Yan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2018 · 2018
Cited alongside, same era.
Efficient sparse-matrix multi-vector product on GPUs. In Proceedings of the 27th International Symposium on High-Performance Parallel and Distributed Computing . ACM, 66–79
Changwan Hong, Aravind Sukumaran-Rajam, Bortik Bandyopadhyay, Jinsung Kim, Süreyya Emre Kurt, Israt Nisa, Shivani Sabhlok, Ümit V Çatalyürek, Srinivasan Parthasarathy, and P Sadayappan. 2018 · 2018
Cited alongside, same era.
Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions
Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen. 2018 · 2018
Cited alongside, same era.
Adaptive sparse tiling for sparse matrix multiplication. In Proceedings of the 24th Symposium on Principles and Practice of Parallel Programming . ACM, 300–314
Changwan Hong, Aravind Sukumaran-Rajam, Israt Nisa, Kunal Singh, and P Sadayappan. 2019 · 2019
Later among the works it cites.
Dissecting the NVidia Turing T4 GPU via microbenchmarking
Zhe Jia, Marco Maggioni, Jeffrey Smith, and Daniele Paolo Scarpazza. 2019 · 2019
Later among the works it cites.
Toward Compact ConvNets via Structure-Sparsity Regularized Filter Pruning
Shaohui Lin, Rongrong Ji, Yuchao Li, Cheng Deng, and Xuelong Li. 2019 · 2019
Later among the works it cites.
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
Mingxing Tan and Quoc V Le. 2019 · 2019
Later among the works it cites.
Triton: an intermediate language and compiler for tiled neural network computations. In Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages . ACM, 10–19
Philippe Tillet, HT Kung, and David Cox. 2019 · 2019
Later among the works it cites.
Structured Pruning of Large Language Models
Ziheng Wang, Jeremy Wohlwend, and Tao Lei. 2019 · 2019
Later among the works it cites.
Balanced sparsity for efficient dnn inference on gpu. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 5676–5683
Zhuliang Yao, Shijie Cao, Wencong Xiao, Chen Zhang, and Lanshun Nie. 2019 · 2019
Later among the works it cites.
Sparse tensor core: Algorithm and hardware co-design for vector-wise sparse neural networks on modern gpus. In Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture . 359–371
Maohua Zhu, Tao Zhang, Zhenyu Gu, and Yuan Xie. 2019 · 2019
Later among the works it cites.
Sparse GPU Kernels for Deep Learning
Trevor Gale, Matei Zaharia, Cliff Young, and Erich Elsen. 2020 · 2020
Closest in time.
A novel data transformation and execution strategy for accelerating sparse matrix multiplication on GPUs. In Proceedings of the 25th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming . 376–388
Peng Jiang, Changwan Hong, and Gagan Agrawal. 2020 · 2020
Closest in time.