Fetching the paper…
Reading the bibliography…
Convolutional neural networks require significant memory bandwidth and storage for intermediate computations, apart from substantial computing resources.
Han, Song, Mao, Huizi, and Dally, William J · 2015
Earlier work this paper cites.
Binarized neural networks
Hubara, I, Courbariaux, M, Soudry, D., El-Yaniv, R., and Bengio, Y · 2016
Earlier work this paper cites.
Xnor-net: Imagenet classification using binary convolutional neural networks
Rastegari, Mohammad, Ordonez, Vicente, Redmon, Joseph, and Farhadi, Ali · 2016
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Zhou, Shuchang, Wu, Yuxin, Ni, Zekun, Zhou, Xinyu, Wen, He, and Zou, Yuheng · 2016
Earlier work this paper cites.
gemmlowp: a small self-contained low-precision gemm library.(2017), 2017
Jacob, Benoit et al · 2017
Earlier work this paper cites.
Towards accurate binary convolutional neural network
Lin, Xiaofan, Zhao, Cong, and Pan, Wei · 2017
Earlier work this paper cites.
8-bit inference with tensorrt
Migacz, S · 2017
Cited alongside, same era.
Pact: Parameterized clipping activation for quantized neural networks
Choi, Jungwook, Wang, Zhuo, Venkataramani, Swagath, Chuang, Pierce I-Jen, Srinivasan, Vijayalakshmi, and Gopalakrishnan, Kailash · 2018
Cited alongside, same era.
Fast adjustable threshold for uniform neural network quantization
Goncharenko, Alexander, Denisov, Andrey, Alyamkin, Sergey, and Terentev, Evgeny · 2018
Cited alongside, same era.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Jacob, Benoit, Kligys, Skirmantas, Chen, Bo, Zhu, Menglong, Tang, Matthew, Howard, Andrew, Adam, Hartwig, and Kalenichenko, Dmitry · 2018
Cited alongside, same era.
Quantizing deep convolutional networks for efficient inference: A whitepaper
Krishnamoorthi, Raghuraman · 2018
Discovering low-precision networks close to full-precision networks for efficient embedded inference
McKinstry, Jeffrey L, Esser, Steven K, Appuswamy, Rathinakumar, Bablani, Deepika, Arthur, John V, Yildiz, Izzet B, and Modha, Dharmendra S · 2018
Closest in time.
Training and inference with integers in deep neural networks
Wu, Shuang, Li, Guoqi, Chen, Feng, and Shi, Luping · 2018
Closest in time.
Low-bit quantization of neural networks for efficient inference
Choukroun, Yoni, Kravchik, Eli, and Kisilev, Pavel · 2019
Closest in time.
Same, same but different-recovering neural network quantization error through weight factorization
Meller, Eldad, Finkelstein, Alexander, Almog, Uri, and Grobman, Mark · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Quantization for rapid deployment of deep neural networks
Lee, Jun Haeng, Ha, Sangwon, Choi, Saerom, Lee, Won-Jo, and Lee, Seungwon · 2018
Cited alongside, same era.
Zhao, Ritchie, Hu, Yuwei, Dotzel, Jordan, De Sa, Christopher, and Zhang, Zhiru · 2019
Closest in time.