Fetching the paper…
Reading the bibliography…
Model compression is instrumental in optimizing deep neural network inference on resource-constrained hardware.
Rao’s distance measure
Colin Atkinson and Ann FS Mitchell · 1981
Earlier work this paper cites.
Skeletonization: A technique for trimming the fat from a network via relevance assessment
Michael C Mozer and Paul Smolensky · 1988
Earlier work this paper cites.
Pruning versus clipping in neural networks
Steven A. Janowsky · 1989
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla · 1989
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David Stork · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context, 2015
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding, 2016
Song Han, Huizi Mao, and William J. Dally · 2016
Earlier work this paper cites.
Information geometry and its applications
Shun-ichi Amari · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text, 2016
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Efficient processing of deep neural networks: A tutorial and survey
Vivienne Sze, Yu-Hsin Chen, Tien-Ju Yang, and Joel S Emer · 2017
Earlier work this paper cites.
Bayesian compression for deep learning, 2017
Christos Louizos, Karen Ullrich, and Max Welling · 2017
Earlier work this paper cites.
More is less: A more complicated network with less inference complexity, 2017
Xuanyi Dong, Junshi Huang, Yi Yang, and Shuicheng Yan · 2017
Earlier work this paper cites.
Mixed precision quantization of convnets via differentiable neural architecture search
Bichen Wu, Yanghan Wang, Peizhao Zhang, Yuandong Tian, Peter Vajda, and Kurt Keutzer · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko · 2018
Earlier work this paper cites.
Snip: Single-shot network pruning based on connection sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip HS Torr · 2018
Earlier work this paper cites.
Faster gaze prediction with dense networks and fisher pruning
Lucas Theis, Iryna Korshunova, Alykhan Tejani, and Ferenc Huszár · 2018
Earlier work this paper cites.
Clip-q: Deep network compression learning by in-parallel pruning-quantization
Frederick Tung and Greg Mori · 2018
Earlier work this paper cites.
Auto-balanced filter pruning for efficient convolutional neural networks
Xiaohan Ding, Guiguang Ding, Jungong Han, and Sheng Tang · 2018
Earlier work this paper cites.
Compressing neural networks using the variational information bottleneck, 2018
Bin Dai, Chen Zhu, and David Wipf · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference, 2018
Adina Williams, Nikita Nangia, and Samuel R. Bowman · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Pact: Parameterized clipping activation for quantized neural networks
Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan · 2018
Earlier work this paper cites.
Hawq: Hessian aware quantization of neural networks with mixed-precision
Zhen Dong, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2019
Earlier work this paper cites.
The state of sparsity in deep neural networks, 2019
Trevor Gale, Erich Elsen, and Sara Hooker · 2019
Earlier work this paper cites.
Importance estimation for neural network pruning
Pavlo Molchanov, Arun Mallya, Stephen Tyree, Iuri Frosio, and Jan Kautz · 2019
Earlier work this paper cites.
Haq: Hardware-aware automated quantization with mixed precision
Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han · 2019
Cited alongside, same era.
Towards optimal structured cnn pruning via generative adversarial learning, 2019
Shaohui Lin, Rongrong Ji, Chenqian Yan, Baochang Zhang, Liujuan Cao, Qixiang Ye, Feiyue Huang, and David Doermann · 2019
Cited alongside, same era.
Mixed precision dnns: All you need is a good parametrization
Stefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama, Javier Alonso Garcia, Stephen Tiedemann, Thomas Kemp, and Akira Nakamura · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
Fast sparse convnets, 2019
Erich Elsen, Marat Dukhan, Trevor Gale, and Karen Simonyan · 2019
Automatic heterogeneous quantization of deep neural networks for low-latency inference on the edge for particle detectors
Claudionor N. Coelho, Aki Kuusela, Shan Li, Hao Zhuang, Jennifer Ngadiuba, Thea Klaeboe Aarrestad, Vladimir Loncar, Maurizio Pierini, Adrian Alan Pol, and Sioni Summers · 2021
Later among the works it cites.
Mobile edge computing enabled 5g health monitoring for internet of medical things: A decentralized game theoretic approach
Zhaolong Ning, Peiran Dong, Xiaojie Wang, Xiping Hu, Lei Guo, Bin Hu, Yi Guo, Tie Qiu, and Ricky Y. K. Kwok · 2021
Later among the works it cites.
How well do sparse imagenet models transfer?
Eugenia Iofinova, Alexandra Peste, Mark Kurtz, and Dan Alistarh · 2021
Later among the works it cites.
Accelerating sparse deep neural networks, 2021
Asit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius · 2021
Later among the works it cites.
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Edgenas: Discovering efficient neural architectures for edge systems
Xiangzhong Luo, Di Liu, Hao Kong, and Weichen Liu · 2020
Cited alongside, same era.
Deep learning in the era of edge computing: Challenges and opportunities, 2020
Mi Zhang, Faen Zhang, Nicholas D. Lane, Yuanchao Shu, Xiao Zeng, Biyi Fang, Shen Yan, and Hui Xu · 2020
Cited alongside, same era.
Inducing and exploiting activation sparsity for fast inference on deep neural networks
Mark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev, John Carr, Michael Goin, William Leiserson, Sage Moore, Bill Nell, Nir Shavit, and Dan Alistarh · 2020
Cited alongside, same era.
What is the state of neural network pruning?
Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag · 2020
Cited alongside, same era.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander Rush · 2020
Cited alongside, same era.
Picking winning tickets before training by preserving gradient flow
Chaoqi Wang, Guodong Zhang, and Roger Grosse · 2020
Cited alongside, same era.
Pruning neural networks without any data by iteratively conserving synaptic flow
Hidenori Tanaka, Daniel Kunin, Daniel L Yamins, and Surya Ganguli · 2020
Cited alongside, same era.
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste · 2021
Later among the works it cites.
A white paper on neural network quantization
Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart van Baalen, and Tijmen Blankevoort · 2021
Later among the works it cites.
Layer-adaptive sparsity for the magnitude-based pruning, 2021
Jaeho Lee, Sejun Park, Sangwoo Mo, Sungsoo Ahn, and Jinwoo Shin · 2021
Later among the works it cites.
M-fac: Efficient matrix-free approximations of second-order information, 2021
Elias Frantar, Eldar Kurtic, and Dan Alistarh · 2021
Later among the works it cites.
Hawq-v3: Dyadic neural network quantization
Zhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami, Jiali Yu, Eric Tan, Leyuan Wang, Qijing Huang, Yida Wang, Michael Mahoney, et al · 2021
Later among the works it cites.
Training deep neural networks with joint quantization and pruning of weights and activations, 2021
Xinyu Zhang, Ian Colbert, Ken Kreutz-Delgado, and Srinjoy Das · 2021
Later among the works it cites.
Uniq: Uniform noise injection for non-uniform quantization of neural networks
Chaim Baskin, Natan Liss, Eli Schwartz, Evgenii Zheltonozhskii, Raja Giryes, Alex M Bronstein, and Avi Mendelson · 2021
Later among the works it cites.
Asks: Convolution with any-shape kernels for efficient neural networks
Guangzhe Liu, Ke Zhang, and Meibo Lv · 2021
Later among the works it cites.
Bcnet: Searching for network width with bilaterally coupled network, 2021
Xiu Su, Shan You, Fei Wang, Chen Qian, Changshui Zhang, and Chang Xu · 2021
Later among the works it cites.
ultralytics/yolov5: v7.0 - yolov5 sota realtime instance segmentation, 2022
Glenn Jocher, Ayush Chaurasia, Alex Stoken, Jirka Borovec, NanoCode012, Yonghye Kwon, Kalen Michael, TaoXie, Jiacong Fang, Imyhxy, , Lorna, Zeng Yifu, Colin Wong, Abhiram V, Diego Montes, Zhiqiang Wang, Cristi Fati, Jebastin Nadar, Laughing, UnglvKitDe, Victor Sonck, Tkianai, YxNONG, Piotr Skalski, Adam Hogan, Dhruv Nair, Max Strobel, and Mrinal Jain · 2022
Later among the works it cites.
Platon: Pruning large transformer models with upper confidence bound of weight importance, 2022
Qingru Zhang, Simiao Zuo, Chen Liang, Alexander Bukharin, Pengcheng He, Weizhu Chen, and Tuo Zhao · 2022
Later among the works it cites.
Ompq: Orthogonal mixed precision quantization, 2022
Yuexiao Ma, Taisong Jin, Xiawu Zheng, Yan Wang, Huixia Li, Yongjian Wu, Guannan Jiang, Wei Zhang, and Rongrong Ji · 2022
Later among the works it cites.
Fit: A metric for model sensitivity
Ben Zandonati, Adrian Alan Pol, Maurizio Pierini, Olya Sirkin, and Tal Kopetz · 2022
Later among the works it cites.
Optimal brain compression: A framework for accurate post-training quantization and pruning
Elias Frantar and Dan Alistarh · 2022
Later among the works it cites.
Opq: Compressing deep neural networks with one-shot pruning-quantization, 2022
Peng Hu, Xi Peng, Hongyuan Zhu, Mohamed M. Sabry Aly, and Jie Lin · 2022
Later among the works it cites.
Computational complexity evaluation of neural network applications in signal processing, 2022
Pedro J. Freire, Sasipim Srivallapanondh, Antonio Napoli, Jaroslaw E. Prilepsky, and Sergei K. Turitsyn · 2022
Later among the works it cites.
Deep model compression based on the training history, 2022
S. H. Shabbeer Basha, Mohammad Farazuddin, Viswanath Pulabaigari, Shiv Ram Dubey, and Snehasis Mukherjee · 2022
Later among the works it cites.
Soks: Automatic searching of the optimal kernel shapes for stripe-wise network pruning
Guangzhe Liu, Ke Zhang, and Meibo Lv · 2022
Later among the works it cites.
Neural network pruning by cooperative coevolution, 2022
Haopu Shang, Jia-Liang Wu, Wenjing Hong, and Chao Qian · 2022
Later among the works it cites.
A fast post-training pruning framework for transformers, 2022
Woosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami · 2022
Later among the works it cites.
Structured pruning is all you need for pruning cnns at initialization, 2022
Yaohui Cai, Weizhe Hua, Hongzheng Chen, G. Edward Suh, Christopher De Sa, and Zhiru Zhang · 2022
Later among the works it cites.
A practical mixed precision algorithm for post-training quantization, 2023
Nilesh Prasad Pandey, Markus Nagel, Mart van Baalen, Yin Huang, Chirag Patel, and Tijmen Blankevoort · 2023
Closest in time.
Mixed precision post training quantization of neural networks with sensitivity guided search, 2023
Clemens JS Schaefer, Elfie Guo, Caitlin Stanton, Xiaofan Zhang, Tom Jablin, Navid Lambert-Shirzad, Jian Li, Chiachen Chou, Siddharth Joshi, and Yu Emma Wang · 2023
Closest in time.
Mixed-precision neural network quantization via learned layer-wise importance, 2023
Chen Tang, Kai Ouyang, Zhi Wang, Yifei Zhu, Yaowei Wang, Wen Ji, and Wenwu Zhu · 2023
Closest in time.
Ma-bert: Towards matrix arithmetic-only bert inference by eliminating complex non-linear functions
Neo Wei Ming, Zhehui Wang, Cheng Liu, Rick Siow Mong Goh, and Tao Luo · 2023
Closest in time.