Y. Cheng, D. Wang, P. Zhou, and T. Zhang, “Model compression and acceleration for deep neural networks: The principles, progress, and challenges,” IEEE Signal Processing Magazine , vol. 35, no. 1, pp. 126–136, 2018
2018
Closest in time.
L. Hou and J. T. Kwok, “Loss-aware weight quantization of deep networks,” in International Conference on Learning Representations , 2018
2018
Closest in time.
P. Gysel, J. Pimentel, M. Motamedi, and S. Ghiasi, “Ristretto: A framework for empirical study of resource-efficient inference in convolutional neural networks,” IEEE Transactions on Neural Networks and Learning Systems , no. 99, pp. 1–6, 2018
2018
Closest in time.
A. Zhou, A. Yao, K. Wang, and Y. Chen, “Explicit loss-error-aware quantization for low-bit deep neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 9426–9435
2018
Closest in time.
P. Yin, S. Zhang, J. Lyu, S. Osher, Y. Qi, and J. Xin, “BinaryRelax: A relaxation approach for training deep neural networks with quantized weights,” SIAM Journal on Imaging Sciences , vol. 11, no. 4, pp. 2205–2223, 2018
2018
Closest in time.
J. Choi, Z. Wang, S. Venkataramani, P. I.-J. Chuang, V. Srinivasan, and K. Gopalakrishnan, “PACT: Parameterized clipping activation for quantized neural networks,” arXiv preprint arXiv:1805.06085 , 2018
Original
2018
Closest in time.
D. Zhang, J. Yang, D. Ye, and G. Hua, “LQ-Nets: Learned quantization for highly accurate and compact deep neural networks,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 365–382
2018
Closest in time.
J. Faraone, N. Fraser, M. Blott, and P. H. Leong, “SYQ: Learning symmetric quantization for efficient deep neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 4300–4309
2018
Closest in time.
X. Zhang, X. Zhou, M. Lin, and J. Sun, “ShuffleNet: An extremely efficient convolutional neural network for mobile devices,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 6848–6856
2018
Closest in time.
F. Tung and G. Mori, “Deep neural network compression by in-parallel pruning-quantization,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2018
2018
Closest in time.
Y. He, J. Lin, Z. Liu, H. Wang, L.-J. Li, and S. Han, “AMC: AutoML for model compression and acceleration on mobile devices,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 784–800
2018
Closest in time.
H. Ren, M. El-Khamy, and J. Lee, “CT-SRCNN: Cascade trained and trimmed deep convolutional neural networks for image super resolution,” in Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV) , 2018
2018
Closest in time.
P. Yin, S. Zhang, J. Lyu, S. Osher, Y. Qi, and J. Xin, “Blended coarse gradient descent for full quantization of deep neural networks,” Research in the Mathematical Sciences , vol. 6, no. 1, p. 14, 2019
2019
Closest in time.
Z.-G. Liu and M. Mattina, “Learning low-precision neural networks without straight-through estimator (STE),” in Proceedings of the International Joint Conference on Artificial Intelligence , 2019, pp. 3066–3072
2019
Closest in time.
J. Yang, X. Shen, J. Xing, X. Tian, H. Li, B. Deng, J. Huang, and X.-s. Hua, “Quantization networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 7308–7316
2019
Closest in time.
R. Gong, X. Liu, S. Jiang, T. Li, P. Hu, J. Lin, F. Yu, and J. Yan, “Differentiable soft quantization: Bridging full-precision and low-bit neural networks,” in Proceedings of the IEEE International Conference on Computer Vision , 2019, pp. 4852–4861
2019
Closest in time.
C. Louizos, M. Reisser, T. Blankevoort, E. Gavves, and M. Welling, “Relaxed quantization for discretized neural networks,” in International Conference on Learning Representations , 2019
2019
Closest in time.
K. Wang, Z. Liu, Y. Lin, J. Lin, and S. Han, “HAQ: Hardware-aware automated quantization with mixed precision,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 8612–8620
2019
Closest in time.
Y. Choi, M. El-Khamy, and J. Lee, “Universal deep neural network compression,” IEEE Journal of Selected Topics in Signal Processing , 2020
2020
Closest in time.
W. Wen, C. Wu, Y. Wang, Y. Chen, and H. Li, “Learning structured sparsity in deep neural networks,” in Advances in Neural Information Processing Systems , 2016, pp. 2074–2082
2082
Closest in time.