Fetching the paper…
Reading the bibliography…
Weight oscillation is an undesirable side effect of quantization-aware training, in which quantized weights frequently jump between two quantized levels, resulting in training instability and a sub-optimal final model.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Convolutional neural networks using logarithmic data representation
Miyashita, D., Lee, E. H., and Murmann, B · 2016
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Zhou, S., Wu, Y., Ni, Z., Zhou, X., Wen, H., and Zou, Y · 2016
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Pact: Parameterized clipping activation for quantized neural networks
Choi, J., Wang, Z., Venkataramani, S., Chuang, P. I.-J., Srinivasan, V., and Gopalakrishnan, K · 2018
Earlier work this paper cites.
Bi-real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm
Liu, Z., Wu, B., Luo, W., Yang, X., Liu, W., and Cheng, K.-T · 2018
Earlier work this paper cites.
Post training 4-bit quantization of convolutional networks for rapid-deployment
Banner, R., Nahshan, Y., and Soudry, D · 2019
Earlier work this paper cites.
Differentiable soft quantization: Bridging full-precision and low-bit neural networks
Gong, R., Liu, X., Jiang, S., Li, T., Hu, P., Lin, J., Yu, F., and Yan, J · 2019
Cited alongside, same era.
Latent weights do not exist: Rethinking binarized neural network optimization
Helwegen, K., Widdicombe, J., Geiger, L., Liu, Z., Cheng, K.-T., and Nusselder, R · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Kenton, J. D. M.-W. C. and Toutanova, L. K · 2019
Cited alongside, same era.
Metapruning: Meta learning for automatic neural network channel pruning
Liu, Z., Mu, H., Zhang, X., Guo, Z., Yang, X., Cheng, K.-T., and Sun, J · 2019
Cited alongside, same era.
Data-free quantization through weight equalization and bias correction
Nagel, M., Baalen, M. v., Blankevoort, T., and Welling, M · 2019
Cited alongside, same era.
Up or down? adaptive rounding for post-training quantization
Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T · 2020
Later among the works it cites.
Xor-net: an efficient computation pipeline for binary neural network inference on edge devices
Zhu, S., Duong, L. H., and Liu, W · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2021
Later among the works it cites.
I-vit: integer-only quantization for efficient vision transformer inference
Li, Z. and Gu, Q · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wu, B., Dai, X., Zhang, P., Wang, Y., Sun, F., Wu, Y., Tian, Y., Vajda, P., Jia, Y., and Keutzer, K · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y · 2019
Cited alongside, same era.
Learned step size quantization
Esser, S. K., McKinstry, J. L., Bablani, D., Appuswamy, R., and Modha, D. S · 2020
Cited alongside, same era.
Additive powers-of-two quantization: An efficient non-uniform discretization for neural networks
Li, Y., Dong, X., and Wang, W · 2020
Cited alongside, same era.
Q-vit: Accurate and fully quantized low-bit vision transformer
Li, Y., Xu, S., Zhang, B., Cao, X., Gao, P., and Guo, G
Cited in the paper.
Q-vit: Fully differentiable quantization for vision transformer
Li, Z., Yang, T., Wang, P., and Cheng, J
Cited in the paper.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B
Cited in the paper.
Later among the works it cites.
Nonuniform-to-uniform quantization: Towards accurate quantization via generalized straight-through estimation
Liu, Z., Cheng, K.-T., Huang, D., Xing, E. P., and Shen, Z · 2022
Later among the works it cites.
Overcoming oscillations in quantization-aware training
Nagel, M., Fournarakis, M., Bondarenko, Y., and Blankevoort, T · 2022
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2022
Later among the works it cites.