Fetching the paper…
Reading the bibliography…
Quantization enables efficient acceleration of deep neural networks by reducing model memory footprint and exploiting low-cost integer math hardware units.
Vector Quantization
Gray, R · 1984
Earlier work this paper cites.
Training Deep Neural Networks with Low Precision Multiplications
Courbariaux, M., Bengio, Y., and David, J.-P · 2014
Earlier work this paper cites.
Compressing Deep Convolutional Networks using Vector Quantization
Gong, Y., Liu, L., Yang, M., and Bourdev, L · 2014
Earlier work this paper cites.
BinaryConnect: Training Deep Neural Networks with Binary Weights during Propagations
Courbariaux, M., Bengio, Y., and David, J.-P · 2015
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2015
Earlier work this paper cites.
Deep Learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Earlier work this paper cites.
Binarized Neural Networks
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y · 2016
Earlier work this paper cites.
Convolutional Neural Networks using Logarithmic Data Representation
Miyashita, D., Lee, E. H., and Murmann, B · 2016
Earlier work this paper cites.
DoReFa-net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
Zhou, S., Wu, Y., Ni, Z., Zhou, X., Wen, H., and Zou, Y · 2016
Earlier work this paper cites.
Zhu, C., Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
Flexpoint: An adaptive Numerical Format for Efficient Training of Deep Neural Networks
Köster, U., Webb, T., Wang, X., Nassar, M., Bansal, A. K., Constable, W., Elibol, O., Gray, S., Hall, S., Hornof, L., et al · 2017
Earlier work this paper cites.
Apprentice: Using Knowledge Distillation Techniques to Improve Low-precision Network Accuracy
Mishra, A. and Marr, D · 2017
Earlier work this paper cites.
Minimum Energy Quantized Neural Networks
Moons, B., Goetschalckx, K., Van Berckelaer, N., and Verhelst, M · 2017
Earlier work this paper cites.
PACT: Parameterized Clipping Activation for Quantized Neural Networks
Choi, J., Wang, Z., Venkataramani, S., Chuang, P. I.-J., Srinivasan, V., and Gopalakrishnan, K · 2018
Cited alongside, same era.
Adaptive Quantization of Neural Networks
Khoram, S. and Li, J · 2018
Cited alongside, same era.
Quantizing Deep Convolutional Networks for Efficient Inference: A Whitepaper
Krishnamoorthi, R · 2018
Cited alongside, same era.
Quantization for Rapid Deployment of Deep Neural Networks
Lee, J. H., Ha, S., Choi, S., Lee, W.-J., and Lee, S · 2018
Cited alongside, same era.
Discovering Low-precision Networks Close to Full-precision Networks for Efficient Embedded Inference
McKinstry, J. L., Esser, S. K., Appuswamy, R., Bablani, D., Arthur, J. V., Yildiz, I. B., and Modha, D. S · 2018
Low Precision Inference on GPUs
Wu, H · 2019
Later among the works it cites.
Zafrir, O., Boudoukh, G., Izsak, P., and Wasserblat, M · 2019
Later among the works it cites.
Improving Neural Network Quantization without Retraining using Outlier Channel Splitting
Zhao, R., Hu, Y., Dotzel, J., De Sa, C., and Zhang, Z · 2019
Later among the works it cites.
ZeroQ: A Novel Zero Shot Quantization Framework
Cai, Y., Yao, Z., Dong, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Later among the works it cites.
Fang, J., Shafiee, A., Abdel-Aziz, H., Thorsley, D., Georgiadis, G., and Hassoun, J · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The NVIDIA Deep Learning Accelerator
Sijstermans, F · 2018
Cited alongside, same era.
Mixed Precision Quantization of ConvNets via Differentiable Neural Architecture Search
Wu, B., Wang, Y., Zhang, P., Tian, Y., Vajda, P., and Keutzer, K · 2018
Cited alongside, same era.
Efficient 8-bit Quantization of Transformer Neural Machine Language Translation Model
Bhandare, A., Sripathi, V., Karkada, D., Menon, V., Choi, S., Datta, K., and Saletore, V · 2019
Cited alongside, same era.
BiScaled-DNN: Quantizing Long-tailed Datastructures with Two Scale Factors for Deep Neural Networks
Jain, S., Venkataramani, S., Srinivasan, V., Choi, J., Gopalakrishnan, K., and Chang, L · 2019
Cited alongside, same era.
Data-free Quantization through Weight Equalization and Bias Correction
Nagel, M., Baalen, M. v., Blankevoort, T., and Welling, M · 2019
Cited alongside, same era.
Fully Quantized Transformer for Improved Translation
Prato, G., Charlaix, E., and Rezagholizadeh, M · 2019
Cited alongside, same era.
MAGNet: A Modular Accelerator Generator for Neural Networks
Venkatesan, R., Shao, Y. S., Wang, M., Clemons, J., Dai, S., Fojtik, M., Keller, B., Klinefelter, A., Pinckney, N. R., Raina, P., et al · 2019
Cited alongside, same era.
NVIDIA A100 Tensor Core GPU Architecture
NVIDIA Corporation · 2020
Later among the works it cites.
Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point
Rouhani, B., Lo, D., Zhao, R., Liu, M., Fowers, J., Ovtcharov, K., Vinogradsky, A., Massengill, S., Yang, L., Bittner, R., et al · 2020
Later among the works it cites.
Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT
Shen, S., Dong, Z., Ye, J., Ma, L., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Later among the works it cites.
And the Bit Goes Down: Revisiting the Quantization of Neural Networks
Stock, P., Joulin, A., Gribonval, R., Graham, B., and Jégou, H · 2020
Later among the works it cites.
Efficient Processing of Deep Neural Networks
Sze, V., Chen, Y.-H., Yang, T.-J., and Emer, J. S · 2020
Later among the works it cites.
Algorithm-Hardware Co-Design of Adaptive Floating-Point Encodings for Resilient Deep Learning Inference
Tambe, T., Yang, E.-Y., Wan, Z., Deng, Y., Reddi, V. J., Rush, A., Brooks, D., and Wei, G.-Y · 2020
Later among the works it cites.
Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
Wu, H., Judd, P., Zhang, X., Isaev, M., and Micikevicius, P · 2020
Later among the works it cites.