Hybrid 8-bit floating point (hfp8) training and inference for deep neural networks
Xiao Sun, Jungwook Choi, Chia-Yu Chen, Naigang Wang, Swagath Venkataramani, Vijayalakshmi (Viji) Srinivasan, Xiaodong Cui, Wei Zhang, and Kailash Gopalakrishnan · 2019
Later among the works it cites.
Lsq+: Improving low-bit quantization through learnable offsets and better initialization
Yash Bhalgat, Jinwon Lee, Markus Nagel, Tijmen Blankevoort, and Nojun Kwak · 2020
Later among the works it cites.
Zeroq: A novel zero shot quantization framework
Original
Yaohui Cai, Zhewei Yao, Zhen Dong, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer · 2020
Later among the works it cites.
Salsanext: Fast, uncertainty-aware semantic segmentation of lidar point clouds for autonomous driving, 2020
Tiago Cortinhal, George Tzelepis, and Eren Erdal Aksoy · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2020
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2020
Later among the works it cites.
Learned step size quantization
Steven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy, and Dharmendra S. Modha · 2020
Later among the works it cites.
Up or down? Adaptive rounding for post-training quantization
Markus Nagel, Rana Ali Amjad, Mart Van Baalen, Christos Louizos, and Tijmen Blankevoort · 2020
Later among the works it cites.
Ultra-low precision 4-bit training of deep neural networks
Xiao Sun, Naigang Wang, Chia-Yu Chen, Jiamin Ni, Ankur Agrawal, Xiaodong Cui, Swagath Venkataramani, Kaoutar El Maghraoui, Vijayalakshmi (Viji) Srinivasan, and Kailash Gopalakrishnan · 2020
Later among the works it cites.
Mixed precision dnns: All you need is a good parametrization
Stefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama, Javier Alonso Garcia, Stephen Tiedemann, Thomas Kemp, and Akira Nakamura · 2020
Later among the works it cites.
A survey of quantization methods for efficient neural network inference, 2021
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer · 2021
Later among the works it cites.
All-you-can-fit 8-bit flexible floating-point format for accurate and memory-efficient inference of deep neural networks, 2021
Cheng-Wei Huang, Tim-Wei Chen, and Juinn-Dar Huang · 2021
Later among the works it cites.
Brecq: Pushing the limit of post-training quantization by block reconstruction
Yuhang Li, Ruihao Gong, Xu Tan, Yang Yang, Peng Hu, Qi Zhang, Fengwei Yu, Wei Wang, and Shi Gu · 2021
Later among the works it cites.
A white paper on neural network quantization
Original
Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart van Baalen, and Tijmen Blankevoort · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
Overcoming oscillations in quantization-aware training
Original
Markus Nagel, Marios Fournarakis, Yelysei Bondarenko, and Tijmen Blankevoort · 2022
Closest in time.