Fetching the paper…
Reading the bibliography…
Post-training quantization attracts increasing attention due to its convenience in deploying quantized neural networks.
Fighting quantization bias with bias
Finkelstein, A.; Almog, U.; and Grobman, M. 2019a · 1906
Earlier work this paper cites.
Fighting Quantization Bias With Bias
Finkelstein, A.; Almog, U.; and Grobman, M. 2019b · 1906
Earlier work this paper cites.
Matrix Analysis
Horn, R. A.; and Johnson, C. R. 1985 · 1985
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Jia, Y.; Shelhamer, E.; Donahue, J.; Karayev, S.; Long, J.; Girshick, R.; Guadarrama, S.; and Darrell, T. 2014 · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Pact: Parameterized clipping activation for quantized neural networks
Choi, J.; Wang, Z.; Venkataramani, S.; Chuang, P. I.-J.; Srinivasan, V.; and Gopalakrishnan, K. 2018 · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Jacob, B.; Kligys, S.; Chen, B.; Zhu, M.; Tang, M.; Howard, A.; Adam, H.; and Kalenichenko, D. 2018 · 2018
Earlier work this paper cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; and Chen, L.-C. 2018 · 2018
Earlier work this paper cites.
ACIQ: Analytical Clipping for Integer Quantization of neural networks
Banner, R.; Nahshan, Y.; Hoffer, E.; and Soudry, D. 2019 · 2019
Earlier work this paper cites.
Post training 4-bit quantization of convolutional networks for rapid-deployment
Banner, R.; Nahshan, Y.; and Soudry, D. 2019 · 2019
Earlier work this paper cites.
Low-bit quantization of neural networks for efficient inference
Choukroun, Y.; Kravchik, E.; Yang, F.; and Kisilev, P. 2019 · 2019
Cited alongside, same era.
LEARNED STEP SIZE QUANTIZATION
Esser, S. K.; McKinstry, J. L.; Bablani, D.; Appuswamy, R.; and Modha, D. S. 2019 · 2019
Cited alongside, same era.
Differentiable soft quantization: Bridging full-precision and low-bit neural networks
Gong, R.; Liu, X.; Jiang, S.; Li, T.; Hu, P.; Lin, J.; Yu, F.; and Yan, J. 2019 · 2019
Cited alongside, same era.
Data-free quantization through weight equalization and bias correction
Nagel, M.; Baalen, M. v.; Blankevoort, T.; and Welling, M. 2019 · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019 · 2019
Cited alongside, same era.
Mnasnet: Platform-aware neural architecture search for mobile
Transformers: State-of-the-Art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Scao, T. L.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. M. 2020 · 2020
Later among the works it cites.
Understanding and Overcoming the Challenges of Efficient Transformer Quantization
Bondarenko, Y.; Nagel, M.; and Blankevoort, T. 2021 · 2021
Later among the works it cites.
Qimera: Data-free Quantization with Synthetic Boundary Supporting Samples
Choi, K.; Hong, D.; Park, N.; Kim, Y.; and Lee, J. 2021 · 2021
Later among the works it cites.
SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian Approximation
Guo, C.; Qiu, Y.; Leng, J.; Gao, X.; Zhang, C.; Liu, Y.; Yang, F.; Zhu, Y.; and Guo, M. 2021 · 2021
Later among the works it cites.
Accurate post training quantization with small calibration sets
Hubara, I.; Nahshan, Y.; Hanani, Y.; Banner, R.; and Soudry, D. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tan, M.; Chen, B.; Pang, R.; Vasudevan, V.; Sandler, M.; Howard, A.; and Le, Q. V. 2019 · 2019
Cited alongside, same era.
Learning channel-wise interactions for binary convolutional neural networks
Wang, Z.; Lu, J.; Tao, C.; Zhou, J.; and Tian, Q. 2019 · 2019
Cited alongside, same era.
Improving neural network quantization without retraining using outlier channel splitting
Zhao, R.; Hu, Y.; Dotzel, J.; De Sa, C.; and Zhang, Z. 2019 · 2019
Cited alongside, same era.
Zeroq: A novel zero shot quantization framework
Cai, Y.; Yao, Z.; Dong, Z.; Gholami, A.; Mahoney, M. W.; and Keutzer, K. 2020 · 2020
Cited alongside, same era.
Up or down? adaptive rounding for post-training quantization
Nagel, M.; Amjad, R. A.; Van Baalen, M.; Louizos, C.; and Blankevoort, T. 2020 · 2020
Cited alongside, same era.
Designing network design spaces
Radosavovic, I.; Kosaraju, R. P.; Girshick, R.; He, K.; and Dollár, P. 2020 · 2020
Cited alongside, same era.
Towards accurate post-training network quantization via bit-split and stitching
Wang, P.; Chen, Q.; He, X.; and Cheng, J. 2020 · 2020
Cited alongside, same era.
BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction
Li, Y.; Gong, R.; Tan, X.; Yang, Y.; Hu, P.; Zhang, Q.; Yu, F.; Wang, W.; and Gu, S. 2021 · 2021
Later among the works it cites.
Post-Training Quantization for Vision Transformer
Liu, Z.; Wang, Y.; Han, K.; Zhang, W.; Ma, S.; and Gao, W. 2021 · 2021
Later among the works it cites.
DNNFusion: accelerating deep neural networks execution with advanced operator fusion
Niu, W.; Guan, J.; Wang, Y.; Agrawal, G.; and Ren, B. 2021 · 2021
Later among the works it cites.
QDrop: Randomly Dropping Quantization for Extremely Low-bit Post-Training Quantization
Wei, X.; Gong, R.; Li, Y.; Liu, X.; and Yu, F. 2022 · 2022
Closest in time.
OPTQ: Accurate Quantization for Generative Pre-trained Transformers
Frantar, E.; Ashkboos, S.; Hoefler, T.; and Alistarh, D. 2023 · 2023
Closest in time.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G.; Lin, J.; Seznec, M.; Wu, H.; Demouth, J.; and Han, S. 2023 · 2023
Closest in time.