Fetching the paper…
Reading the bibliography…
Recently, the idea of using FP8 as a number format for neural network training has been floating around the deep learning world.
High-level area estimation
Buyuksahin, K. M. and Najm, F · 2002
Earlier work this paper cites.
Microsoft COCO: common objects in context
Lin, T., Maire, M., Belongie, S. J., Bourdev, L. D., Girshick, R. B., Hays, J., Perona, P., Ramanan, D., Doll’a r, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
The pascal visual object classes challenge: A retrospective
Everingham, M., Eslami, S., Van Gool, L., Williams, C., Winn, J., and Zisserman, A · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Gupta, S., Agrawal, A., Gopalakrishnan, K., and Narayanan, P · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Modified fused multiply and add for exact low precision product accumulation
Brunie, N · 2017
Earlier work this paper cites.
Rethinking atrous convolution for semantic image segmentation, 2017
Chen, L.-C., Papandreou, G., Schroff, F., and Adam, H · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2018
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Quantizing deep convolutional networks for efficient inference: A whitepaper
Krishnamoorthi, R · 2018
Earlier work this paper cites.
URL https://github.com/tonylins/pytorch-mobilenet-v2
Lins, T., 2018 · 2018
Earlier work this paper cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Earlier work this paper cites.
Training deep neural networks with 8-bit floating point numbers
Wang, N., Choi, J., Brand, D., Chen, C.-Y., and Gopalakrishnan, K · 2018
Earlier work this paper cites.
Deeplabv3 code and model, 2018
Zhang, J · 2018
Earlier work this paper cites.
Semantickitti: A dataset for semantic scene understanding of lidar sequences
Behley, J., Garbade, M., Milioto, A., Quenzel, J., Behnke, S., Stachniss, C., and Gall, J · 2019
Earlier work this paper cites.
Searching for mobilenetv3
Howard, A., Sandler, M., Chu, G., Chen, L.-C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., Le, Q. V., and Adam, H · 2019
Earlier work this paper cites.
Data-free quantization through weight equalization and bias correction
Nagel, M., van Baalen, M., Blankevoort, T., and Welling, M · 2019
Cited alongside, same era.
Nvidia: Apex automatic mixed precision, 2019
Nvidia · 2019
Cited alongside, same era.
Hybrid 8-bit floating point (hfp8) training and inference for deep neural networks
Sun, X., Choi, J., Chen, C.-Y., Wang, N., Venkataramani, S., Srinivasan, V. V., Cui, X., Zhang, W., and Gopalakrishnan, K · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q. V · 2019
Cited alongside, same era.
Training high-performance and large-scale deep neural networks with full 8-bit integers
Yanga, Y., Wua, S., Dengb, L., Yanc, T., Xieb, Y., and Li, G · 2019
Cited alongside, same era.
Q8bert: Quantized 8bit bert
Zafrir, O., Boudoukh, G., Izsak, P., and Wasserblat, M · 2019
Understanding deep learning (still) requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2021
Later among the works it cites.
Nvidia hopper architecture in-depth, 2022
Andersch, M., Palmer, G., Krashinsky, R., Stam, N., Mehta, V., Brito, G., and Ramaswamy, S · 2022
Later among the works it cites.
The case for 4-bit precision: k-bit inference scaling laws
Dettmers, T. and Zettlemoyer, L · 2022
Later among the works it cites.
Llm.int8(): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Later among the works it cites.
ultralytics/yolov5: v7.0 - YOLOv5 SOTA Realtime Instance Segmentation, November 2022
Jocher, G., Chaurasia, A., Stoken, A., Borovec, J., NanoCode012, Kwon, Y., Michael, K., TaoXie, Fang, J., imyhxy, Lorna, Yifu, Z., Wong, C., V, A., Montes, D., Wang, Z., Fati, C., Nadar, J., Laughing, UnglvKitDe, Sonck, V., tkianai, yxNONG, Skalski, P., Hogan, A., Nair, D., Strobel, M., and Jain, M · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Lsq+: Improving low-bit quantization through learnable offsets and better initialization
Bhalgat, Y., Lee, J., Nagel, M., Blankevoort, T., and Kwak, N · 2020
Cited alongside, same era.
nuscenes: A multimodal dataset for autonomous driving
Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O · 2020
Cited alongside, same era.
Learned step size quantization
Esser, S. K., McKinstry, J. L., Bablani, D., Appuswamy, R., and Modha, D. S · 2020
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Cited alongside, same era.
Understanding and overcoming the challenges of efficient transformer quantization
Bondarenko, Y., Nagel, M., and Blankevoort, T · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2021
Cited alongside, same era.
Fp8 quantization: The power of the exponent
Kuzmin, A., van Baalen, M., Ren, Y., Nagel, M., Peters, J., and Blankevoort, T · 2022
Later among the works it cites.
Efficientformer: Vision transformers at mobilenet speed
Li, Y., Yuan, G., Wen, Y., Hu, E., Evangelidis, G., Tulyakov, S., Wang, Y., and Ren, J · 2022
Later among the works it cites.
Simple and efficient architectures for semantic segmentation
Mehta, D., Skliar, A., Ben Yahia, H., Borse, S., Porikli, F., Habibian, A., and Blankevoort, T · 2022
Later among the works it cites.
Micikevicius, P., Stosic, D., Judd, P., Kamalu, J., Oberman, S., Shoeybi, M., Siu, M., Wu, H., Burgess, N., Ha, S., Grisenthwaite, R., Mellempudi, N., Cornea, M., Heinecke, A., and Dubey, P · 2022
Later among the works it cites.
Overcoming oscillations in quantization-aware training
Nagel, M., Fournarakis, M., Bondarenko, Y., and Blankevoort, T · 2022
Later among the works it cites.
Nvidia: Transformer engine, 2022
Nvidia · 2022
Later among the works it cites.
Neural network quantization with ai model efficiency toolkit (aimet)
Siddegowda, S., Fournarakis, M., Nagel, M., Blankevoort, T., Patel, C., and Khobare, A · 2022
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2022
Later among the works it cites.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Yao, Z., Aminabadi, R. Y., Zhang, M., Wu, X., Li, C., and He, Y · 2022
Later among the works it cites.
Gptq: Accurate quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2023
Closest in time.
8-bit numerical formats for deep neural networks
Noune, B., Jones, P., Justus, D., Masters, D., and Luschi, C · 2023
Closest in time.
Shared microexponents: A little shifting goes a long way
Rouhani, B., Zhao, R., Elango, V., Shafipour, R., Hall, M., Mesmakhosroshahi, M., More, A., Melnick, L., Varatkar, M. G. G., Shao, L., Kolhe, G., Melts, D., Klar, J., L’Heureux, R., Perry, M., Burger, D., Chung, E., Deng, Z., Naghshineh, S., Park, J., and Naumov, M · 2023
Closest in time.