Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Learning structured sparsity in deep neural networks
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2016
Cited alongside, same era.
Coordinating filters for faster deep neural networks
Wei Wen, Cong Xu, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
Incremental network quantization: Towards lossless cnns with low-precision weights
Original
Aojun Zhou, Anbang Yao, Yiwen Guo, Lin Xu, and Yurong Chen · 2017
Cited alongside, same era.
Pact: Parameterized clipping activation for quantized neural networks
Original
Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan · 2018
Cited alongside, same era.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko · 2018
Cited alongside, same era.
Value-aware quantization for training and inference of neural networks
Eunhyeok Park, Sungjoo Yoo, and Peter Vajda · 2018
Cited alongside, same era.
Model compression via distillation and quantization
Original
A. Polino, R. Pascanu, and D. Alistarh · 2018
Cited alongside, same era.
An analytical method to determine minimum per-layer precision of deep neural networks
Charbel Sakr and Naresh Shanbhag · 2018
Cited alongside, same era.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Original
Song Han, Huizi Mao, and William J Dally
Cited in the paper.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally
Cited in the paper.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun
Cited in the paper.