Fetching the paper…
Reading the bibliography…
Recent advances in deep learning methods such as LLMs and Diffusion models have created a need for improved quantization methods that can meet the computational demands of these modern architectures while maintaining accuracy.
Improving neural network quantization without retraining using outlier channel splitting
Zhao, R., Hu, Y., Dotzel, J., Sa, C. D., and Zhang, Z · 1901
Earlier work this paper cites.
Mixed Precision Training With 8-bit Floating Point, May 2019
Mellempudi, N., Srinivasan, S., Das, D., and Kaul, B · 1905
Earlier work this paper cites.
Fang, J., Shafiee, A., Abdel-Aziz, H., Thorsley, D., Georgiadis, G., and Hassoun, J · 2002
Earlier work this paper cites.
LSQ+: improving low-bit quantization through learnable offsets and better initialization
Bhalgat, Y., Lee, J., Nagel, M., Blankevoort, T., and Kwak, N · 2004
Earlier work this paper cites.
Improving the speed of neural networks on cpus
Vanhoucke, V., Senior, A., and Mao, M. Z · 2011
Earlier work this paper cites.
Microsoft COCO: common objects in context
Lin, T., Maire, M., Belongie, S. J., Bourdev, L. D., Girshick, R. B., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Resiliency of deep neural networks under quantization
Sung, W., Shin, S., and Hwang, K · 2015
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Convolutional neural networks using logarithmic data representation
Miyashita, D., Lee, E. H., and Murmann, B · 2016
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Zhou, S., Ni, Z., Zhou, X., Wen, H., Wu, Y., and Zou, Y · 2016
Earlier work this paper cites.
Deep learning with low precision by half-wave gaussian quantization
Cai, Z., He, X., Sun, J., and Vasconcelos, N · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Klambauer, G., and Hochreiter, S · 2017
Earlier work this paper cites.
Ternary neural networks with fine-grained quantization
Mellempudi, N., Kundu, A., Mudigere, D., Das, D., Kaul, B., and Dubey, P · 2017
Earlier work this paper cites.
8-bit inference with tensorrt, 2017
Migacz, S · 2017
Earlier work this paper cites.
Carvana image masking challenge, 2017
Shaler, B., DanGill, Maggie, McDonald, M., Patricia, and Cukierski, W · 2017
Earlier work this paper cites.
Incremental network quantization: Towards lossless CNNs with low-precision weights
Zhou, A., Yao, A., Guo, Y., Xu, L., and Chen, Y · 2017
Cited alongside, same era.
PACT: parameterized clipping activation for quantized neural networks
Choi, J., Wang, Z., Venkataramani, S., Chuang, P. I., Srinivasan, V., and Gopalakrishnan, K · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., and Kalenichenko, D · 2018
Cited alongside, same era.
Marian: Fast neural machine translation in C++
Junczys-Dowmunt, M., Grundkiewicz, R., Dwojak, T., Hoang, H., Heafield, K., Neckermann, T., Seide, F., Germann, U., Fikri Aji, A., Bogoychev, N., Martins, A. F. T., and Birch, A · 2018
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
Learned step size quantization
Esser, S. K., McKinstry, J. L., Bablani, D., Appuswamy, R., and Modha, D. S · 2020
Later among the works it cites.
Additive powers-of-two quantization: An efficient non-uniform discretization for neural networks
Li, Y., Dong, X., and Wang, W · 2020
Later among the works it cites.
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
Zhang, J., Zhao, Y., Saleh, M., and Liu, P · 2020
Later among the works it cites.
A survey of quantization methods for efficient neural network inference
Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Quantizing deep convolutional networks for efficient inference: A whitepaper
Krishnamoorthi, R · 2018
Cited alongside, same era.
Yolov3: An incremental improvement
Redmon, J. and Farhadi, A · 2018
Cited alongside, same era.
Training Deep Neural Networks with 8-bit Floating Point Numbers
Wang, N., Choi, J., Brand, D., Chen, C.-Y., and Gopalakrishnan, K · 2018
Cited alongside, same era.
Efficient 8-bit quantization of transformer neural machine language translation model
Bhandare, A., Sripathi, V., Karkada, D., Menon, V., Choi, S., Datta, K., and Saletore, V · 2019
Cited alongside, same era.
Low-bit quantization of neural networks for efficient inference
Choukroun, Y., Kravchik, E., Yang, F., and Kisilev, P · 2019
Cited alongside, same era.
Deep learning recommendation model for personalization and recommendation systems
Naumov, M., Mudigere, D., Shi, H.-J. M., Huang, J., Sundaraman, N., Park, J., Wang, X., Gupta, U., Wu, C.-J., Azzolini, A. G., et al · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A · 2021
Later among the works it cites.
Efficient quantization techniques for deep neural networks
Jiang, C · 2021
Later among the works it cites.
I-bert: Integer-only bert quantization
Kim, S., Gholami, A., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Later among the works it cites.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Later among the works it cites.
FP8 Quantization: The Power of the Exponent, August 2022
Kuzmin, A., Van Baalen, M., Ren, Y., Nagel, M., Peters, J., and Blankevoort, T · 2022
Later among the works it cites.
Biogpt: generative pre-trained transformer for biomedical text generation and mining
Luo, R., Sun, L., Xia, Y., Qin, T., Zhang, S., Poon, H., and Liu, T.-Y · 2022
Later among the works it cites.
Micikevicius, P., Stosic, D., Burgess, N., Cornea, M., Dubey, P., Grisenthwaite, R., Ha, S., Heinecke, A., Judd, P., Kamalu, J., et al · 2022
Later among the works it cites.
8-bit Numerical Formats for Deep Neural Networks, June 2022
Noune, B., Jones, P., Justus, D., Masters, D., and Luschi, C · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilić, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., et al · 2022
Later among the works it cites.
Outlier suppression: Pushing the limit of low-bit transformer language models
Wei, X., Zhang, Y., Zhang, X., Gong, R., Zhang, S., Zhang, Q., Yu, F., and Liu, X · 2022
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Demouth, J., and Han, S · 2022
Later among the works it cites.
Code llama: Open foundation models for code, 2023
Rozière, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., Kozhevnikov, A., Evtimov, I., Bitton, J., Bhatt, M., Ferrer, C. C., Grattafiori, A., Xiong, W., Défossez, A., Copet, J., Azhar, F., Touvron, H., Martin, L., Usunier, N., Scialom, T., and Synnaeve, G · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Closest in time.