Fetching the paper…
Reading the bibliography…
The last few years have seen gigantic leaps in algorithms and systems to support efficient deep learning inference.
Sparse tiling for stationary iterative methods
Michelle Mills Strout, Larry Carter, Jeanne Ferrante, and Barbara Kreaseck. 2004 · 2004
Earlier work this paper cites.
High-performance implementation of the level-3 BLAS
Kazushige Goto and Robert Van De Geijn. 2008 · 2008
Earlier work this paper cites.
The University of Florida sparse matrix collection
Timothy A Davis and Yifan Hu. 2011 · 2011
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014 · 2014
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally. 2015 · 2015
Earlier work this paper cites.
EIE: efficient inference engine on compressed deep neural network. In 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) . IEEE, 243–254
Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A Horowitz, and William J Dally. 2016 · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
Faster cnns with direct sparse convolutions and guided pruning
Jongsoo Park, Sheng Li, Wei Wen, Ping Tak Peter Tang, Hai Li, Yiran Chen, and Pradeep Dubey. 2016 · 2016
Earlier work this paper cites.
Gpu kernels for block-sparse weights
Scott Gray, Alec Radford, and Diederik P Kingma. 2017 · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017 · 2017
Earlier work this paper cites.
Cutlass: Fast linear algebra in cuda c++
Andrew Kerr, Duane Merrill, Julien Demouth, and John Tran. 2017 · 2017
Earlier work this paper cites.
Learning Sparse Neural Networks through L _ 0 L\_0 Regularization
Christos Louizos, Max Welling, and Diederik P Kingma. 2017 · 2017
Earlier work this paper cites.
Variational dropout sparsifies deep neural networks. In Proceedings of the 34th International Conference on Machine Learning-Volume 70 . JMLR. org, 2498–2507
Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov. 2017 · 2017
Earlier work this paper cites.
TVM: end-to-end optimization stack for deep learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Haichen Shen, Eddie Q Yan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2018 · 2018
Cited alongside, same era.
Artificial-intelligence hardware: New opportunities for semiconductor companies
Siddarth Madhav Andrea Queirolo Gaurav Batra, Zach Jacobson and Nick Santhanam. 2018 · 2018
Cited alongside, same era.
OpenVINO deep learning workbench: Comprehensive analysis and tuning of neural networks inference. In Proceedings of the IEEE International Conference on Computer Vision Workshops . 0–0
Yury Gorbachev, Mikhail Fedorov, Iliya Slavutin, Artyom Tugarev, Marat Fatekhov, and Yaroslav Tarkan. 2019 · 2019
Later among the works it cites.
Adaptive sparse tiling for sparse matrix multiplication. In Proceedings of the 24th Symposium on Principles and Practice of Parallel Programming . ACM, 300–314
Changwan Hong, Aravind Sukumaran-Rajam, Israt Nisa, Kunal Singh, and P Sadayappan. 2019 · 2019
Later among the works it cites.
Optimizing { \{ CNN } \} Model Inference on CPUs. In 2019 { \{ USENIX } \} Annual Technical Conference ( { \{ USENIX } \} { \{ ATC } \} 19) . 1025–1040
Yizhi Liu, Yao Wang, Ruofei Yu, Mu Li, Vin Sharma, and Yida Wang. 2019 · 2019
Later among the works it cites.
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
Mingxing Tan and Quoc V Le. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Changwan Hong, Aravind Sukumaran-Rajam, Bortik Bandyopadhyay, Jinsung Kim, Süreyya Emre Kurt, Israt Nisa, Shivani Sabhlok, Ümit V Çatalyürek, Srinivasan Parthasarathy, and P Sadayappan. 2018 · 2018
Cited alongside, same era.
TensorRT
Nvidia. 2018 · 2018
Cited alongside, same era.
Design principles for sparse matrix multiplication on the GPU. In European Conference on Parallel Processing . Springer, 672–687
Carl Yang, Aydın Buluç, and John D Owens. 2018 · 2018
Cited alongside, same era.
Escoin: Efficient Sparse Convolutional Neural Network Inference on GPUs
Xuhao Chen. 2019 · 2019
Cited alongside, same era.
A closer look at structured pruning for neural network compression
Elliot Crowley, Jack Turner, Amos Storkey, and O’Boyle Michael. 2019 · 2019
Cited alongside, same era.
Erich Elsen, Marat Dukhan, Trevor Gale, and Karen Simonyan. 2019 · 2019
Cited alongside, same era.
Rigging the Lottery: Making All Tickets Winners
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen. 2019 · 2019
Cited alongside, same era.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker. 2019 · 2019
Cited alongside, same era.
Triton: an intermediate language and compiler for tiled neural network computations. In Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages . ACM, 10–19
Philippe Tillet, HT Kung, and David Cox. 2019 · 2019
Later among the works it cites.
Structured Pruning of Large Language Models
Ziheng Wang, Jeremy Wohlwend, and Tao Lei. 2019 · 2019
Later among the works it cites.
Balanced sparsity for efficient dnn inference on gpu. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 5676–5683
Zhuliang Yao, Shijie Cao, Wencong Xiao, Chen Zhang, and Lanshun Nie. 2019 · 2019
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Sparse GPU Kernels for Deep Learning
Trevor Gale, Matei Zaharia, Cliff Young, and Erich Elsen. 2020 · 2020
Later among the works it cites.
MNN: A Universal and Efficient Inference Engine
Xiaotang Jiang, Huan Wang, Yiliu Chen, Ziqi Wu, Lichuan Wang, Bin Zou, Yafeng Yang, Zongyang Cui, Yu Cai, Tianhang Yu, et al · 2020
Later among the works it cites.
Neural network compression framework for fast model inference
Alexander Kozlov, Ivan Lazarevich, Vasily Shamporov, Nikolay Lyalyushkin, and Yury Gorbachev. 2020 · 2020
Later among the works it cites.
Movement Pruning: Adaptive Sparsity by Fine-Tuning
Victor Sanh, Thomas Wolf, and Alexander M Rush. 2020 · 2020
Later among the works it cites.
SparseRT: Accelerating Unstructured Sparsity on GPUs for Deep Learning Inference
Ziheng Wang. 2020 · 2020
Later among the works it cites.