Fetching the paper…
Reading the bibliography…
Many neural network quantization techniques have been developed to decrease the computational and memory footprint of deep learning.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L. Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, N. Léonard, and Aaron C. Courville · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (tripod): the tripod statement
Gary S. Collins, Johannes B. Reitsma, Douglas G. Altman, and Karel GM Moons · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Fixed point quantization of deep convolutional networks
Darryl Lin, Sachin Talathi, and Sreekanth Annapureddy · 2016
Earlier work this paper cites.
Convolutional neural networks using logarithmic data representation
Daisuke Miyashita, Edward H. Lee, and Boris Murmann · 2016
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman · 2018
Earlier work this paper cites.
PACT: parameterized clipping activation for quantized neural networks
Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan · 2018
Earlier work this paper cites.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna M. Wallach, Hal Daumé III, and Kate Crawford · 2018
Earlier work this paper cites.
The dataset nutrition label: A framework to drive higher data quality standards
Sarah Holland, Ahmed Hosny, Sarah Newman, Joshua Joseph, and Kasia Chmielinski · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko · 2018
Earlier work this paper cites.
Quantizing deep convolutional networks for efficient inference: A whitepaper
Raghuraman Krishnamoorthi · 2018
Earlier work this paper cites.
Bi-real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm
Zechun Liu, Baoyuan Wu, Wenhan Luo, Xin Yang, Wei Liu, and Kwang-Ting Cheng · 2018
Earlier work this paper cites.
Discovering low-precision networks close to full-precision networks for efficient embedded inference
Jeffrey L. McKinstry, Steven K. Esser, Rathinakumar Appuswamy, Deepika Bablani, John V. Arthur, Izzet B. Yildiz, and Dharmendra S. Modha · 2018
Earlier work this paper cites.
Precision highway for ultra low-precision quantization
Eunhyeok Park, Dongyoung Kim, Sungjoo Yoo, and Peter Vajda · 2018
Cited alongside, same era.
Mixed precision quantization of convnets via differentiable neural architecture search
Bichen Wu, Yanghan Wang, Peizhao Zhang, Yuandong Tian, Peter Vajda, and Kurt Keutzer · 2018
Cited alongside, same era.
Lq-nets: Learned quantization for highly accurate and compact deep neural networks
Dongqing Zhang, Jiaolong Yang, Dongqiangzi Ye, and Gang Hua · 2018
Cited alongside, same era.
Factsheets: Increasing trust in ai services through supplier’s declarations of conformity
M. Arnold, R. K. E. Bellamy, M. Hind, S. Houde, S. Mehta, A. Mojsilović, R. Nair, K. N. Ramamurthy, A. Olteanu, D. Piorkowski, D. Reimer, J. Richards, J. Tsay, and K. R. Varshney · 2019
Cited alongside, same era.
Post training 4-bit quantization of convolutional networks for rapid-deployment
Ron Banner, Yury Nahshan, and Daniel Soudry · 2019
Once-for-all: Train one network and specialize it for efficient deployment
Han Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang, and Song Han · 2020
Later among the works it cites.
Zeroq: A novel zero shot quantization framework
Yaohui Cai, Zhewei Yao, Zhen Dong, Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2020
Later among the works it cites.
HAWQ-V2: hessian aware trace-weighted quantization of neural networks
Zhen Dong, Zhewei Yao, Yaohui Cai, Daiyaan Arfeen, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer · 2020
Later among the works it cites.
Learned step size quantization
Steven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy, and Dharmendra S. Modha · 2020
Later among the works it cites.
Improving post training neural quantization: Layer-wise calibration and integer programming, 2020
Itay Hubara, Yury Nahshan, Yair Hanani, Ron Banner, and Daniel Soudry · 2020
Later among the works it cites.
Up or down? adaptive rounding for post-training quantization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Low-bit quantization of neural networks for efficient inference
Y. Choukroun, E. Kravchik, F. Yang, and P. Kisilev · 2019
Cited alongside, same era.
Hawq: Hessian aware quantization of neural networks with mixed-precision
Zhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer · 2019
Cited alongside, same era.
Fully quantized network for object detection
R. Li, Y. Wang, F. Liang, H. Qin, J. Yan, and R. Fan · 2019
Cited alongside, same era.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Cited alongside, same era.
Data-free quantization through weight equalization and bias correction
M. Nagel, M. V. Baalen, T. Blankevoort, and M. Welling · 2019
Cited alongside, same era.
Loss aware post-training quantization
Yury Nahshan, Brian Chmiel, Chaim Baskin, Evgenii Zheltonozhskii, Ron Banner, Alexander M. Bronstein, and Avi Mendelson · 2019
Cited alongside, same era.
Haq: Hardware-aware automated quantization with mixed precision
Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, and Song Han · 2019
Cited alongside, same era.
Markus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos, and Tijmen Blankevoort · 2020
Later among the works it cites.
Mixed precision dnns: All you need is a good parametrization
S. Uhlich, L. Mauch, F. Cardinaux, K. Yoshiyama, J. A. Garcia, S. Tiedemann, T. Kemp, and A. Nakamura · 2020
Later among the works it cites.
Towards accurate post-training network quantization via bit-split and stitching
Peisong Wang, Qiang Chen, Xiangyu He, and Jian Cheng · 2020
Later among the works it cites.
Generative low-bitwidth data free quantization
Shoukai Xu, Haokun Li, Bohan Zhuang, Jing Liu, Jiezhang Cao, Chuangrun Liang, and Mingkui Tan · 2020
Later among the works it cites.
Hawqv3: Dyadic neural network quantization, 2020
Zhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami, Jiali Yu, Eric Tan, Leyuan Wang, Qijing Huang, Yida Wang, Michael W. Mahoney, and Kurt Keutzer · 2020
Later among the works it cites.
TernaryBERT: Distillation-aware ultra-low bit BERT
Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu · 2020
Later among the works it cites.
I-bert: Integer-only bert quantization, 2021
Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer · 2021
Closest in time.
Post-training quantization with multiple points: Mixed precision without mixed precision
Xingchao Liu, M. Ye, Dengyong Zhou, and Qiang Liu · 2021
Closest in time.
Fracbits: Mixed precision quantization via fractional bit-widths, 2021
Linjie Yang and Qing Jin · 2021
Closest in time.