Fetching the paper…
Reading the bibliography…
Neural network pruning and quantization techniques are almost as old as neural networks themselves.
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla · 1989
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
Babak Hassibi, David G Stork, and Gregory J Wolff · 1993
Earlier work this paper cites.
The pascal visual object classes (voc) challenge
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2010
Earlier work this paper cites.
A Mathematical Introduction to Compressive Sensing
Simon Foucart and Holger Rauhut · 2013
Earlier work this paper cites.
CVX: Matlab software for disciplined convex programming, version 2.1
Michael Grant and Stephen Boyd · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Ellipsoid bounds for convex quadratic integer programming
Christoph Buchheim, Ruth Huebner, and Anita Schoebel · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Accelerating very deep convolutional networks for classification and detection
Xiangyu Zhang, Jianhua Zou, Kaiming He, and Jian Sun · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Fixed point quantization of deep convolutional networks
Darryl Lin, Sachin Talathi, and Sreekanth Annapureddy · 2016
Earlier work this paper cites.
Fixed point quantization of deep convolutional networks
Darryl D. Lin, Sachin S. Talathi, and V. Sreekanth Annapureddy · 2016
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou, Zekun Ni, Xinyu Zhou, He Wen, Yuxin Wu, and Yuheng Zou · 2016
Earlier work this paper cites.
Rethinking atrous convolution for semantic image segmentation, 2017
Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam · 2017
Earlier work this paper cites.
Channel pruning for accelerating very deep neural networks
Yihui He, Xiangyu Zhang, and Jian Sun · 2017
Earlier work this paper cites.
Learning sparse neural networks through l _ 0 l\_0 regularization
Christos Louizos, Max Welling, and Diederik P Kingma · 2017
Earlier work this paper cites.
Exploring sparsity in recurrent neural networks
Sharan Narang, Erich Elsen, Gregory Diamos, and Shubho Sengupta · 2017
Earlier work this paper cites.
Mixed-integer quadratic programming is in np
Alberto Del Pia, Santanu S Dey, and Marco Molinaro · 2017
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Michael Zhu and Suyog Gupta · 2017
Earlier work this paper cites.
PACT: parameterized clipping activation for quantized neural networks
Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko · 2018
Cited alongside, same era.
Snip: Single-shot network pruning based on connection sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip HS Torr · 2018
Cited alongside, same era.
Learning sparse neural networks through l 0 l_{0} regularization
Christos Louizos, Max Welling, and Diederik P Kingma · 2018
Cited alongside, same era.
A semidefinite programming method for integer convex quadratic minimization
Jaehyun Park and Stephen Boyd · 2018
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Cited alongside, same era.
Characterising bias in compressed models
Sara Hooker, Nyalleng Moorosi, Gregory Clark, Samy Bengio, and Emily Denton · 2020
Later among the works it cites.
Up or down? Adaptive rounding for post-training quantization
Markus Nagel, Rana Ali Amjad, Mart Van Baalen, Christos Louizos, and Tijmen Blankevoort · 2020
Later among the works it cites.
Efficientdet: Scalable and efficient object detection, 2020
Mingxing Tan, Ruoming Pang, and Quoc V. Le · 2020
Later among the works it cites.
Mixed precision dnns: All you need is a good parametrization
Stefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama, Javier Alonso Garcia, Stephen Tiedemann, Thomas Kemp, and Akira Nakamura · 2020
Later among the works it cites.
Bayesian bits: Unifying quantization and pruning
Mart van Baalen, Christos Louizos, Markus Nagel, Rana Ali Amjad, Ying Wang, Tijmen Blankevoort, and Max Welling · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep neural network compression by in-parallel pruning-quantization
Frederick Tung and Greg Mori · 2018
Cited alongside, same era.
Nisp: Pruning networks using neuron importance score propagation
Ruichi Yu, Ang Li, Chun-Fu Chen, Jui-Hsin Lai, Vlad I Morariu, Xintong Han, Mingfei Gao, Ching-Yung Lin, and Larry S Davis · 2018
Cited alongside, same era.
Post training 4-bit quantization of convolutional networks for rapid-deployment
Ron Banner, Yury Nahshan, and Daniel Soudry · 2019
Cited alongside, same era.
Low-bit quantization of neural networks for efficient inference
Yoni Choukroun, Eli Kravchik, and Pavel Kisilev · 2019
Cited alongside, same era.
Hawq: Hessian aware quantization of neural networks with mixed-precision
Zhen Dong, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2019
Cited alongside, same era.
Stabilizing the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M Roy, and Michael Carbin · 2019
Cited alongside, same era.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker · 2019
Cited alongside, same era.
Automatic neural network compression by sparsity-quantization joint learning: A constrained optimization-based approach
Haichuan Yang, Shupeng Gui, Yuhao Zhu, and Ji Liu · 2020
Later among the works it cites.
Opq: Compressing deep neural networks with one-shot pruning-quantization
Peng Hu, Xi Peng, Hongyuan Zhu, Mohamed M Sabry Aly, and Jie Lin · 2021
Later among the works it cites.
Accurate post training quantization with small calibration sets
Itay Hubara, Yury Nahshan, Yair Hanani, Ron Banner, and Daniel Soudry · 2021
Later among the works it cites.
An empirical comparison of quantization, pruning and low-rank neural network compression using the lc toolkit
Yerlan Idelbayev and Miguel Á Carreira-Perpiñán · 2021
Later among the works it cites.
Brecq: Pushing the limit of post-training quantization by block reconstruction
Yuhang Li, Ruihao Gong, Xu Tan, Yang Yang, Peng Hu, Qi Zhang, Fengwei Yu, Wei Wang, and Shi Gu · 2021
Later among the works it cites.
A white paper on neural network quantization
Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart van Baalen, and Tijmen Blankevoort · 2021
Later among the works it cites.
A white paper on neural network quantization
Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart van Baalen, and Tijmen Blankevoort · 2021
Later among the works it cites.
BLOOM (revision 4ab0472), 2022
BigScience Workshop · 2022
Later among the works it cites.
Optimal brain compression: A framework for accurate post-training quantization and pruning
Elias Frantar and Dan Alistarh · 2022
Later among the works it cites.
Fp8 quantization: The power of the exponent
Andrey Kuzmin, Mart Van Baalen, Yuwei Ren, Markus Nagel, Jorn Peters, and Tijmen Blankevoort · 2022
Later among the works it cites.
Accelerating sparse deep neural networks
Asit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius · 2022
Later among the works it cites.
Overcoming oscillations in quantization-aware training
Markus Nagel, Marios Fournarakis, Yelysei Bondarenko, and Tijmen Blankevoort · 2022
Later among the works it cites.
Cyclical pruning for sparse neural networks
Suraj Srinivas, Andrey Kuzmin, Markus Nagel, Mart van Baalen, Andrii Skliar, and Tijmen Blankevoort · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Sparsegpt: Massive language models can be accurately pruned in one-shot, 2023
Elias Frantar and Dan Alistarh · 2023
Closest in time.
Openllama: An open reproduction of llama, May 2023
Xinyang Geng and Hao Liu · 2023
Closest in time.
Gurobi Optimizer Reference Manual, 2023
Gurobi Optimization, LLC · 2023
Closest in time.
The cost of compression: Investigating the impact of compression on parametric knowledge in language models
Satya Sai Srinath Namburi, Makesh Sreedhar, Srinath Srinivasan, and Frederic Sala · 2023
Closest in time.