Fetching the paper…
Reading the bibliography…
Neural network quantization is frequently used to optimize model size, latency and power consumption for on-device deployment of neural networks.
Addressing for random-access storage
W. W. Peterson · 1957
Earlier work this paper cites.
The art of computer programming
D. E. Knuth · 1997
Earlier work this paper cites.
The pascal visual object classes challenge: A retrospective
M. Everingham, S. Eslami, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Rethinking atrous convolution for semantic image segmentation
L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko · 2018
Earlier work this paper cites.
Quantizing deep convolutional networks for efficient inference: A whitepaper
R. Krishnamoorthi · 2018
Earlier work this paper cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2018
Earlier work this paper cites.
Mixed precision quantization of convnets via differentiable neural architecture search
B. Wu, Y. Wang, P. Zhang, Y. Tian, P. Vajda, and K. Keutzer · 2018
Cited alongside, same era.
Post training 4-bit quantization of convolutional networks for rapid-deployment
R. Banner, Y. Nahshan, and D. Soudry · 2019
Cited alongside, same era.
Hawq: Hessian aware quantization of neural networks with mixed-precision
Z. Dong, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer · 2019
Cited alongside, same era.
Learned step size quantization
S. K. Esser, J. L. McKinstry, D. Bablani, R. Appuswamy, and D. S. Modha · 2019
Cited alongside, same era.
Searching for mobilenetv3
A. Howard, M. Sandler, G. Chu, L.-C. Chen, B. Chen, M. Tan, W. Wang, Y. Zhu, R. Pang, V. Vasudevan, et al · 2019
Cited alongside, same era.
Bayesian bits: Unifying quantization and pruning
M. Van Baalen, C. Louizos, M. Nagel, R. A. Amjad, Y. Wang, T. Blankevoort, and M. Welling · 2020
Later among the works it cites.
Ompq: Orthogonal mixed precision quantization
Y. Ma, T. Jin, X. Zheng, Y. Wang, H. Li, G. Jiang, W. Zhang, and R. Ji · 2021
Later among the works it cites.
A white paper on neural network quantization
M. Nagel, M. Fournarakis, R. A. Amjad, Y. Bondarenko, M. van Baalen, and T. Blankevoort · 2021
Later among the works it cites.
Fracbits: Mixed precision quantization via fractional bit-widths
L. Yang and Q. Jin · 2021
Later among the works it cites.
Sdq: Stochastic differentiable quantization with mixed precision
X. Huang, Z. Shen, S. Li, Z. Liu, H. Xianghong, J. Wicaksana, E. Xing, and K.-T. Cheng · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Data-free quantization through weight equalization and bias correction
M. Nagel, M. v. Baalen, T. Blankevoort, and M. Welling · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
M. Tan and Q. Le · 2019
Cited alongside, same era.
Haq: Hardware-aware automated quantization with mixed precision
K. Wang, Z. Liu, Y. Lin, J. Lin, and S. Han · 2019
Cited alongside, same era.
Lsq+: Improving low-bit quantization through learnable offsets and better initialization
Y. Bhalgat, J. Lee, M. Nagel, T. Blankevoort, and N. Kwak · 2020
Cited alongside, same era.
Up or down? adaptive rounding for post-training quantization
M. Nagel, R. A. Amjad, M. Van Baalen, C. Louizos, and T. Blankevoort · 2020
Cited alongside, same era.
M. Nagel, M. Fournarakis, Y. Bondarenko, and T. Blankevoort · 2022
Later among the works it cites.
Power-of-two quantization for low bitwidth and hardware compliant neural networks
D. Przewlocka-Rus, S. S. Sarwar, H. E. Sumbul, Y. Li, and B. De Salvo · 2022
Later among the works it cites.
Neural network quantization with ai model efficiency toolkit (aimet)
S. Siddegowda, M. Fournarakis, M. Nagel, T. Blankevoort, C. Patel, and A. Khobare · 2022
Later among the works it cites.
Simulated quantization, real power savings
M. van Baalen, B. Kahne, E. Mahurin, A. Kuzmin, A. Skliar, M. Nagel, and T. Blankevoort · 2022
Later among the works it cites.
Fit: A metric for model sensitivity
B. Zandonati, A. A. Pol, M. Pierini, O. Sirkin, and T. Kopetz · 2022
Later among the works it cites.