Mixed precision quantization of convnets via differentiable neural architecture search
Original
Bichen Wu, Yanghan Wang, Peizhao Zhang, Yuandong Tian, Peter Vajda, and Kurt Keutzer · 2018
Later among the works it cites.
LQ-Nets: Learned quantization for highly accurate and compact deep neural networks
Dongqing Zhang, Jiaolong Yang, Dongqiangzi Ye, and Gang Hua · 2018
Later among the works it cites.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun · 2018
Later among the works it cites.
Adaptive quantization for deep neural network
Yiren Zhou, Seyed-Mohsen Moosavi-Dezfooli, Ngai-Man Cheung, and Pascal Frossard · 2018
Later among the works it cites.
https://github.com/amirgholami/zeroq.git, Dec. 2019
2019
Later among the works it cites.
Cat: Compression-aware training for bandwidth reduction
Original
Chaim Baskin, Brian Chmiel, Evgenii Zheltonozhskii, Ron Banner, Alex M Bronstein, and Avi Mendelson · 2019
Later among the works it cites.
Hawq-v2: Hessian aware trace-weighted quantization of neural networks
Original
Zhen Dong, Zhewei Yao, Yaohui Cai, Daiyaan Arfeen, Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2019
Later among the works it cites.
Hawq: Hessian aware quantization of neural networks with mixed-precision
Zhen Dong, Zhewei Yao, Amir Gholami, Michael Mahoney, and Kurt Keutzer · 2019
Later among the works it cites.
The knowledge within: Methods for data-free model compression
Matan Haroush, Itay Hubara, Elad Hoffer, and Daniel Soudry · 2019
Later among the works it cites.
Low-bit quantization of neural networks for efficient inference
Eli Kravchik, Fan Yang, Pavel Kisilev, and Yoni Choukroun · 2019
Later among the works it cites.
Fully quantized network for object detection
Rundong Li, Yan Wang, Feng Liang, Hongwei Qin, Junjie Yan, and Rui Fan · 2019
Later among the works it cites.
Same, same but different-recovering neural network quantization error through weight factorization
Original
Eldad Meller, Alexander Finkelstein, Uri Almog, and Mark Grobman · 2019
Later among the works it cites.
Data-free quantization through weight equalization and bias correction
Markus Nagel, Mart van Baalen, Tijmen Blankevoort, and Max Welling · 2019
Later among the works it cites.
Q-bert: Hessian based ultra low precision quantization of bert
Original
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2019
Later among the works it cites.
HAQ: Hardware-aware automated quantization
Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han · 2019
Later among the works it cites.
Improving neural network quantization without retraining using outlier channel splitting
Ritchie Zhao, Yuwei Hu, Jordan Dotzel, Chris De Sa, and Zhiru Zhang · 2019
Later among the works it cites.