Fetching the paper…
Reading the bibliography…
Neural network quantization has become an important research area due to its great impact on deployment of large models on resource constrained devices.
Optimal brain damage
Yann LeCun, John S Denker, and Sara A Solla · 1990
Earlier work this paper cites.
Joe Staines and David Barber · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Earlier work this paper cites.
Binaryconnect: Training deep neural networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
Fengfu Li, Bo Zhang, and Bin Liu · 2016
Earlier work this paper cites.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J Maddison, Andriy Mnih, and Yee Whye Teh · 2016
Cited alongside, same era.
Xnor-net: Imagenet classification using binary convolutional neural networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi · 2016
Cited alongside, same era.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou · 2016
Cited alongside, same era.
Chenzhuo Zhu, Song Han, Huizi Mao, and William J Dally · 2016
Cited alongside, same era.
Deep learning with low precision by half-wave gaussian quantization
Variational network quantization
Jan Achterhold, Jan Mathias Koehler, Anke Schmeink, and Tim Genewein · 2018
Closest in time.
Uniq: Uniform noise injection for the quantization of neural networks
Chaim Baskin, Eli Schwartz, Evgenii Zheltonozhskii, Natan Liss, Raja Giryes, Alex M Bronstein, and Avi Mendelson · 2018
Closest in time.
Syq: Learning symmetric quantization for efficient deep neural networks
Julian Faraone, Nicholas Fraser, Michaela Blott, and Philip HW Leong · 2018
Closest in time.
Ristretto: A framework for empirical study of resource-efficient inference in convolutional neural networks
Philipp Gysel, Jon Pimentel, Mohammad Motamedi, and Soheil Ghiasi · 2018
Closest in time.
Low-precision floating-point schemes for neural network training
Marc Ortiz, Adrián Cristal, Eduard Ayguadé, and Marc Casas · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhaowei Cai, Xiaodong He, Jian Sun, and Nuno Vasconcelos · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Cited alongside, same era.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko · 2017
Cited alongside, same era.
Deep convolutional neural network inference with floating-point weights and fixed-point activations
Liangzhen Lai, Naveen Suda, and Vikas Chandra · 2017
Cited alongside, same era.
Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy
Asit Mishra and Debbie Marr · 2017
Cited alongside, same era.
Wrpn: wide reduced-precision networks
Asit Mishra, Eriko Nurvitadhi, Jeffrey J Cook, and Debbie Marr · 2017
Cited alongside, same era.
Soft weight-sharing for neural network compression
Karen Ullrich, Edward Meeds, and Max Welling · 2017
Cited alongside, same era.
Incremental network quantization: Towards lossless cnns with low-precision weights
Aojun Zhou, Anbang Yao, Yiwen Guo, Lin Xu, and Yurong Chen · 2017
Cited alongside, same era.
Probabilistic binary neural networks
Jorn WT Peters and Max Welling · 2018
Closest in time.
Model compression via distillation and quantization
Antonio Polino, Razvan Pascanu, and Dan Alistarh · 2018
Closest in time.
Learning discrete weights using the local reparameterization trick
Oran Shayer, Dan Levi, and Ethan Fetaya · 2018
Closest in time.
A quantization-friendly separable convolution for mobilenets
Tao Sheng, Chen Feng, Shaojie Zhuo, Xiaopeng Zhang, Liang Shen, and Mickey Aleksic · 2018
Closest in time.
Training and inference with integers in deep neural networks
Shuang Wu, Guoqi Li, Feng Chen, and Luping Shi · 2018
Closest in time.
Explicit loss-error-aware quantization for low-bit deep neural networks
Aojun Zhou, Anbang Yao, Kuan Wang, and Yurong Chen · 2018
Closest in time.