Fetching the paper…
Reading the bibliography…
4-bit and lower precision mobile models are required due to the ever-increasing demand for better energy efficiency in mobile devices.
Deng, J., Dong, W., Socher, R., Li, L., Li, K., Li, F.: Imagenet: A large-scale hierarchical image database. Computer Vision and Pattern Recognition (CVPR) (2009)
2009
Earlier work this paper cites.
Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. Neural Information Processing Systems (NeurIPS) (2012)
2012
Earlier work this paper cites.
2015
Earlier work this paper cites.
Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition. International Conference on Learning Representations (ICLR) (2015)
2015
Earlier work this paper cites.
Albericio, J., Judd, P., Hetherington, T.H., Aamodt, T.M., Jerger, N.D.E., Moshovos, A.: Cnvlutin: Ineffectual-neuron-free deep neural network computing. International Symposium on Computer Architecture (ISCA) (2016)
2016
Earlier work this paper cites.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. Computer Vision and Pattern Recognition (CVPR) (2016)
2016
Earlier work this paper cites.
Rastegari, M., Ordonez, V., Redmon, J., Farhadi, A.: Xnor-net: Imagenet classification using binary convolutional neural networks. European Conference on Computer Vision (ECCV) (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., Bengio, Y.: Quantized neural networks: Training neural networks with low precision weights and activations. Journal of Machine Learning Research (JMLR) (2017)
2017
Earlier work this paper cites.
Loshchilov, I., Hutter, F.: SGDR: stochastic gradient descent with warm restarts. International Conference on Learning Representations (ICLR) (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Zhou, S., Wang, Y., Wen, H., He, Q., Zou, Y.: Balanced quantization: An effective and efficient approach to quantized neural networks. Journal of Computer Science and Technology (2017)
2017
Earlier work this paper cites.
Zhu, C., Han, S., Mao, H., Dally, W.J.: Trained ternary quantization. INternational Conference on Learning Representations (ICLR) (2017)
2017
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
Howard, A., Sandler, M., Chu, G., Chen, L., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., Le, Q.V., Adam, H.: Searching for mobilenetv3. International Conference on Computer Vision (ICCV) (2019)
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D.P., Wilson, A.G.: Averaging weights leads to wider optima and better generalization. Uncertainty in Artificial Intelligence (UAI) (2018)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Liu, Z., Wu, B., Luo, W., Yang, X., Liu, W., Cheng, K.: Bi-real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm. European Conference on Computer Vision (ECCV) (2018)
2018
Cited alongside, same era.
Mishra, A.K., Nurvitadhi, E., Cook, J.J., Marr, D.: WRPN: wide reduced-precision networks. International Conference on Learning Representations (ICLR) (2018)
2018
Cited alongside, same era.
Park, E., Kim, D., Yoo, S.: Energy-efficient neural network accelerator based on outlier-aware low-precision computation. International Symposium on Computer Architecture (ISCA) (2018)
2018
Cited alongside, same era.
Sandler, M., Howard, A.G., Zhu, M., Zhmoginov, A., Chen, L.: Mobilenetv2: Inverted residuals and linear bottlenecks. Computer Vision and Pattern Recognition (CVPR) (2018)
2018
Cited alongside, same era.
Sharma, H., Park, J., Suda, N., Lai, L., Chau, B., Kim, J.K., Chandra, V., Esmaeilzadeh, H.: Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural networks. International Symposium on Computer Architecture (ISCA) (2018)
2018
Cited alongside, same era.
Tulloch, A., Jia, Y.: Quantization and training of neural networks for efficient integer-arithmetic-only inference. Conference on Computer Vision and Pattern Recognition (CVPR) (2018)
2018
Cited alongside, same era.
Jung, S., Son, C., Lee, S., Son, J., Han, J., Kwak, Y., Hwang, S.J., Choi, C.: Learning to quantize deep networks by optimizing quantization intervals with task loss. Computer Vision and Pattern Recognition (CVPR) (2019)
2019
Later among the works it cites.
2019
Later among the works it cites.
Song, J., Cho, Y., Park, J.S., Jang, J.W., Lee, S., Song, J.H., Lee, J.G., Kang, I.: 7.1 an 11.5 tops/w 1024-mac butterfly structure dual-core sparsity-aware neural processing unit in 8nm flagship mobile soc. International Solid-State Circuits Conference (ISSCC) (2019)
2019
Later among the works it cites.
Tan, M., Chen, B., Pang, R., Vasudevan, V., Sandler, M., Howard, A., Le, Q.V.: Mnasnet: Platform-aware neural architecture search for mobile. Computer Vision and Pattern Recognition (CVPR) (2019)
2019
Later among the works it cites.
Wang, T., Xiong, J., Xu, X., Shi, Y.: SCNN: A general distribution based statistical convolutional neural network with application to video object detection. Association for the Advancement of Artificial Intelligence (AAAI) (2019)
2019
Later among the works it cites.
Wu, H.: NVIDIA Low Precision Inference on GPU. GPU Technology Conference (2019)
2019
Later among the works it cites.
Krizhevsky, A., Nair, V., Hinton, G.: The cifar-10 dataset. https://www.cs.toronto.edu/~kriz/cifar.html
2020
Closest in time.
Int4 precision for ai inference. https://devblogs.nvidia.com/int4-for-ai-inference/
2020
Closest in time.
Samsung low-power npu solution for ai deep learning. https://news.samsung.com/global/samsung-electronics-introduces-a-high-speed-low-power-npu-solution-for-ai-deep-learning
2020
Closest in time.
Snapdragon neural processing engine sdk. https://developer.qualcomm.com/docs/snpe/index.html
2020
Closest in time.