Fetching the paper…
Reading the bibliography…
We present an overview of techniques for quantizing convolutional neural networks for inference with integer weights and activations.
B. Polyak, “New stochastic approximation type procedures,” Jan 1990
1990
Earlier work this paper cites.
M. Courbariaux, Y. Bengio, and J. David, “Binaryconnect: Training deep neural networks with binary weights during propagations,” 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the Inception Architecture for Computer Vision,” Dec. 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” 2015
2015
Earlier work this paper cites.
Software available from tensorflow.org
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015 · 2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” 2015
2015
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the Knowledge in a Neural Network,” Mar. 2015
2015
Earlier work this paper cites.
2016
Cited alongside, same era.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “EIE: efficient inference engine on compressed deep neural network,” 2016
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” 2016
2016
Cited alongside, same era.
A. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications,” Apr. 2017
2017
Cited alongside, same era.
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference,” Dec. 2017
A. K. Mishra and D. Marr, “Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy,” 2017
2017
Later among the works it cites.
A. K. Mishra, E. Nurvitadhi, J. J. Cook, and D. Marr, “WRPN: wide reduced-precision networks,” CoRR
2017
Later among the works it cites.
A. K. Mishra and D. Marr, “Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy,” 2017
2017
Later among the works it cites.
M. Sandler, A. G. Howard, M. Zhu, A. Zhmoginov, and L. Chen, “Inverted residuals and linear bottlenecks: Mobile networks for classification, detection and segmentation,” 2018
2018
Closest in time.
https://www.qualcomm.com/news/onq/2018/02/01/how-can-snapdragon-845s-new-ai-boost-your-smartphones-iq
Qualcomm onQ blog, “How can Snapdragon 845’s new AI boost your smartphone’s IQ?.” · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
http://on-demand.gputechconf.com/gtc/2017/presentation/s7310-8-bit-inference-with-tensorrt.pdf
Nvidia, “8 bit inference with TensorRT.” · 2017
Cited alongside, same era.
2017
Cited alongside, same era.
B. Zoph, V. Vasudevan, J. Shlens, and Q. V. Le, “Learning transferable architectures for scalable image recognition,” 2017
2017
Cited alongside, same era.
https://github.com/google/gemmlowp
GEMMLOWP, “Gemmlowp: a small self-contained low-precision GEMM library.”
Cited in the paper.
https://intel.github.io/mkl-dnn/index.html
Intel(R) MKL-DNN, “Intel(R) Math Kernel Library for Deep Neural Networks.”
Cited in the paper.
http://arm-software.github.io/CMSIS_5/NN/html/index.html
ARM, “Arm cmsis nn software library.”
Cited in the paper.
http://nvdla.org/
Nvidia, “The nvidia deep learning accelerator.”
Cited in the paper.
Closest in time.
A. Polino, R. Pascanu, and D. Alistarh, “Model compression via distillation and quantization,” 2018
2018
Closest in time.
T. Sheng, C. Feng, S. Zhuo, X. Zhang, L. Shen, and M. Aleksic, “A quantization-friendly separable convolution for mobilenets,” 2018
2018
Closest in time.