Fetching the paper…
Reading the bibliography…
Quantization is an effective method for reducing memory footprint and inference time of Neural Networks, e.g., for efficient inference in the cloud, especially at the edge.
Experimental determination of precision requirements for back-propagation training of artificial neural networks
Krste Asanovic and Nelson Morgan · 1991
Earlier work this paper cites.
Some large-scale matrix computation problems
Zhaojun Bai, Gark Fahey, and Gene Golub · 1996
Earlier work this paper cites.
Randomized algorithms for estimating the trace of an implicit symmetric positive semi-definite matrix
Haim Avron and Sivan Toledo · 2011
Earlier work this paper cites.
Randomized algorithms for matrices and data
M. W. Mahoney · 2011
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
BinaryConnect: Training deep neural networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and William J Dally · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Xnor-net: Imagenet classification using binary convolutional neural networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Quantized convolutional neural networks for mobile devices
Jiaxiang Wu, Cong Leng, Yuhang Wang, Qinghao Hu, and Jian Cheng · 2016
Cited alongside, same era.
DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou · 2016
Cited alongside, same era.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár · 2017
Cited alongside, same era.
Incremental network quantization: Towards lossless CNNs with low-precision weights
Aojun Zhou, Anbang Yao, Yiwen Guo, Lin Xu, and Yurong Chen · 2017
Cited alongside, same era.
Value-aware quantization for training and inference of neural networks
Eunhyeok Park, Sungjoo Yoo, and Peter Vajda · 2018
Later among the works it cites.
Mixed precision quantization of convnets via differentiable neural architecture search
Bichen Wu, Yanghan Wang, Peizhao Zhang, Yuandong Tian, Peter Vajda, and Kurt Keutzer · 2018
Later among the works it cites.
Large batch size training of neural networks with adversarial training and second-order information
Zhewei Yao, Amir Gholami, Kurt Keutzer, and Michael W. Mahoney · 2018
Later among the works it cites.
Hessian-based analysis of large batch training and robustness to adversaries
Zhewei Yao, Amir Gholami, Qi Lei, Kurt Keutzer, and Michael W. Mahoney · 2018
Later among the works it cites.
LQ-Nets: Learned quantization for highly accurate and compact deep neural networks
Dongqing Zhang, Jiaolong Yang, Dongqiangzi Ye, and Gang Hua · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan · 2018
Cited alongside, same era.
SqueezeNext: Hardware-aware neural network design
Amir Gholami, Kiseok Kwon, Bichen Wu, Zizheng Tai, Xiangyu Yue, Peter Jin, Sicheng Zhao, and Kurt Keutzer · 2018
Cited alongside, same era.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko · 2018
Cited alongside, same era.
Quantizing deep convolutional networks for efficient inference: A whitepaper
Raghuraman Krishnamoorthi · 2018
Cited alongside, same era.
The Mathematics of Data
M. W. Mahoney, J. C. Duchi, and A. C. Gilbert, editors · 2018
Cited alongside, same era.
Later among the works it cites.
HAWQ: Hessian aware quantization of neural networks with mixed-precision
Zhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer · 2019
Closest in time.
Fully quantized network for object detection
Rundong Li, Yan Wang, Feng Liang, Hongwei Qin, Junjie Yan, and Rui Fan · 2019
Closest in time.
Q-BERT: Hessian based ultra low precision quantization of bert
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer · 2019
Closest in time.
HAQ: Hardware-aware automated quantization
Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han · 2019
Closest in time.
Trust region based adversarial attack on neural networks
Zhewei Yao, Amir Gholami, Peng Xu, Kurt Keutzer, and Michael W. Mahoney · 2019
Closest in time.