Fetching the paper…
Reading the bibliography…
When training neural networks with simulated quantization, we observe that quantized weights can, rather unexpectedly, oscillate between two grid-points.
Optimization by simulated annealing
Kirkpatrick, S., Gelatt, C., and Vecchi, M · 1983
Earlier work this paper cites.
Neural networks for machine learning, lectures 15b
Hinton, G · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
1.1 computing’s energy problem (and what we can do about it)
Horowitz, M · 2014
Earlier work this paper cites.
Deep learning with limited numerical precision
Gupta, S., Agrawal, A., Gopalakrishnan, K., and Narayanan, P · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Zhou, S., Ni, Z., Zhou, X., Wen, H., Wu, Y., and Zou, Y · 2016
Earlier work this paper cites.
Deep learning with low precision by half-wave gaussian quantization
Cai, Z., He, X., Sun, J., and Vasconcelos, N · 2017
Earlier work this paper cites.
PACT: parameterized clipping activation for quantized neural networks
Choi, J., Wang, Z., Venkataramani, S., Chuang, P. I., Srinivasan, V., and Gopalakrishnan, K · 2018
Earlier work this paper cites.
Quantizing deep convolutional networks for efficient inference: A whitepaper
Krishnamoorthi, R · 2018
Earlier work this paper cites.
Probabilistic binary neural networks
Peters, J. W. T. and Welling, M · 2018
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S · 2018
Earlier work this paper cites.
Post training 4-bit quantization of convolutional networks for rapid-deployment
Banner, R., Nahshan, Y., and Soudry, D · 2019
Cited alongside, same era.
Metaquant: Learning to quantize by learning to penetrate non-differentiable quantization
Chen, S., Wang, W., and Pan, S. J · 2019
Cited alongside, same era.
Differentiable soft quantization: Bridging full-precision and low-bit neural networks
Gong, R., Liu, X., Jiang, S., Li, T., Hu, P., Lin, J., Yu, F., and Yan, J · 2019
Cited alongside, same era.
Latent weights do not exist: Rethinking binarized neural network optimization
Helwegen, K., Widdicombe, J., Geiger, L., Liu, Z., Cheng, K.-T., and Nusselder, R · 2019
Cited alongside, same era.
Jain, S. R., Gural, A., Wu, M., and Dick, C · 2019
Cited alongside, same era.
Loss aware post-training quantization
Nahshan, Y., Chmiel, B., Baskin, C., Zheltonozhskii, E., Banner, R., Bronstein, A. M., and Mendelson, A · 2020
Later among the works it cites.
Quantization aware training with absolute-cosine regularization for automatic speech recognition
Nguyen, H. D., Alexandridis, A., and Mouchtaris, A · 2020
Later among the works it cites.
PROFIT: A novel training method for sub-4-bit mobilenet models
Park, E. and Yoo, S · 2020
Later among the works it cites.
Low bias low variance gradient estimates for hierarchical boolean stochastic networks
Pervez, A., Cohen, T., and Gavves, E · 2020
Later among the works it cites.
Mixed precision dnns: All you need is a good parametrization
Uhlich, S., Mauch, L., Cardinaux, F., Yoshiyama, K., Garcia, J. A., Tiedemann, S., Kemp, T., and Nakamura, A · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Relaxed quantization for discretized neural networks
Louizos, C., Reisser, M., Blankevoort, T., Gavves, E., and Welling, M · 2019
Cited alongside, same era.
Data-free quantization through weight equalization and bias correction
Nagel, M., van Baalen, M., Blankevoort, T., and Welling, M · 2019
Cited alongside, same era.
Quantization networks
Yang, J., Shen, X., Xing, J., Tian, X., Li, H., Deng, B., Huang, J., and Hua, X.-s · 2019
Cited alongside, same era.
Understanding straight-through estimator in training activation quantized neural nets
Yin, P., Lyu, J., Zhang, S., Osher, S. J., Qi, Y., and Xin, J · 2019
Cited alongside, same era.
Lsq+: Improving low-bit quantization through learnable offsets and better initialization
Bhalgat, Y., Lee, J., Nagel, M., Blankevoort, T., and Kwak, N · 2020
Cited alongside, same era.
Learned step size quantization
Esser, S. K., McKinstry, J. L., Bablani, D., Appuswamy, R., and Modha, D. S · 2020
Cited alongside, same era.
Position-based scaled gradient for model quantization and pruning
Kim, J., Yoo, K., and Kwak, N · 2020
Cited alongside, same era.
Quantization-guided training for compact tiny{ml} models
Chai, S. M · 2021
Later among the works it cites.
Differentiable model compression via pseudo quantization noise
Défossez, A., Adi, Y., and Synnaeve, G · 2021
Later among the works it cites.
Training with quantization noise for extreme model compression
Fan, A., Stock, P., Graham, B., Grave, E., Gribonval, R., Jegou, H., and Joulin, A · 2021
Later among the works it cites.
Improving low-precision network quantization via bin regularization
Han, T., Li, D., Liu, J., Tian, L., and Shan, Y · 2021
Later among the works it cites.
Network quantization with element-wise gradient scaling
J. Lee, D. Kim, B. H · 2021
Later among the works it cites.
A white paper on neural network quantization
Nagel, M., Fournarakis, M., Amjad, R. A., Bondarenko, Y., van Baalen, M., and Blankevoort, T · 2021
Later among the works it cites.
Logarithmic unbiased quantization: Practical 4-bit training in deep learning
Chmiel, B., Banner, R., Hoffer, E., Yaacov, H. B., and Soudry, D · 2022
Closest in time.