Fetching the paper…
Reading the bibliography…
When quantizing neural networks, assigning each floating-point weight to its nearest fixed-point value is the predominant approach.
“neural” computation of decisions in optimization problems
Hopfield, J. J. and Tank, D. W · 1985
Earlier work this paper cites.
A vlsi architecture for high-performance, low-cost, on-chip learning
Hammerstrom, D · 1990
Earlier work this paper cites.
Learning with limited numerical precision using the cascade-correlation algorithm
Hoehfeld, M. and Fahlman, S. E · 1992
Earlier work this paper cites.
Finite precision error analysis of neural network hardware implementations
Holi, J. L. and Hwang, J. N · 1993
Earlier work this paper cites.
The cross-entropy method for combinatorial and continuous optimization
Rubinstein, R · 1999
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
The unconstrained binary quadratic programming problem: a survey
Kochenberger, G., Hao, J.-K., Glover, F., Lewis, M., Lü, Z., Wang, H., and Wang, Y · 2014
Earlier work this paper cites.
The pascal visual object classes challenge: A retrospective
Everingham, M., Eslami, S., Van Gool, L., Williams, C., Winn, J., and Zisserman, A · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Gupta, S., Agrawal, A., Gopalakrishnan, K., and Narayanan, P · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Fixed point quantization of deep convolutional networks
Lin, D. D., Talathi, S. S., and Annapureddy, V. S · 2016
Earlier work this paper cites.
Accelerating very deep convolutional networks for classification and detection
Zhang, X., Zou, J., He, K., and Sun, J · 2016
Cited alongside, same era.
Practical gauss-newton optimisation for deep learning
Botev, A., Ritter, H., and Barber, D · 2017
Cited alongside, same era.
Channel pruning for accelerating very deep neural networks
He, Y., Zhang, X., and Sun, J · 2017
Cited alongside, same era.
Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy
Mishra, A. K. and Marr, D · 2017
Cited alongside, same era.
Encoder-decoder with atrous separable convolution for semantic image segmentation
Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H · 2018
Cited alongside, same era.
PACT: parameterized clipping activation for quantized neural networks
HAWQ: hessian aware quantization of neural networks with mixed-precision
Dong, Z., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2019
Later among the works it cites.
Fighting quantization bias with bias
Finkelstein, A., Almog, U., and Grobman, M · 2019
Later among the works it cites.
Jain, S. R., Gural, A., Wu, M., and Dick, C · 2019
Later among the works it cites.
QKD: quantization-aware knowledge distillation
Kim, J., Bhalgat, Y., Lee, J., Patel, C., and Kwak, N · 2019
Later among the works it cites.
Relaxed quantization for discretized neural networks
Louizos, C., Reisser, M., Blankevoort, T., Gavves, E., and Welling, M · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Choi, J., Wang, Z., Venkataramani, S., Chuang, P. I., Srinivasan, V., and Gopalakrishnan, K · 2018
Cited alongside, same era.
A survey on methods and theories of quantized neural networks
Guo, Y · 2018
Cited alongside, same era.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., and Kalenichenko, D · 2018
Cited alongside, same era.
Quantizing deep convolutional networks for efficient inference: A whitepaper
Krishnamoorthi, R · 2018
Cited alongside, same era.
Learning sparse neural networks through l 0 l_{0} regularization
Louizos, C., Welling, M., and Kingma, D. P · 2018
Cited alongside, same era.
Two-step quantization for low-bit neural networks
Wang, P., Hu, Q., Zhang, Y., Zhang, C., Liu, Y., and Cheng, J · 2018
Cited alongside, same era.
Post training 4-bit quantization of convolutional networks for rapid-deployment
Banner, R., Nahshan, Y., and Soudry, D · 2019
Cited alongside, same era.
Data-free quantization through weight equalization and bias correction
Nagel, M., van Baalen, M., Blankevoort, T., and Welling, M · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Later among the works it cites.
Improving neural network quantization without retraining using outlier channel splitting
Zhao, R., Hu, Y., Dotzel, J., Sa, C. D., and Zhang, Z · 2019
Later among the works it cites.
Zeroq: A novel zero shot quantization framework
Cai, Y., Yao, Z., Dong, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Closest in time.
Learned step size quantization
Esser, S. K., McKinstry, J. L., Bablani, D., Appuswamy, R., and Modha, D. S · 2020
Closest in time.
And the bit goes down: Revisiting the quantization of neural networks
Stock, P., Joulin, A., Gribonval, R., Graham, B., and Jégou, H · 2020
Closest in time.
Mixed precision dnns: All you need is a good parametrization
Uhlich, S., Mauch, L., Yoshiyama, K., Cardinaux, F., García, J. A., Tiedemann, S., Kemp, T., and Nakamura, A · 2020
Closest in time.