Fetching the paper…
Reading the bibliography…
Quantization is a key technique to reduce the resource requirement and improve the performance of neural network deployment.
Optimal Brain Damage
Le Cun, Yann Le Cun, John S Denker, and Sara A Sol · 1990
Earlier work this paper cites.
LLVM: A compilation framework for lifelong program analysis & transformation
Chris Lattner and Vikram Adve · 2004
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1
Matthieu Courbariaux, Itay Hubara, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding
Song Han, Huizi Mao, and William J. Dally · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Rethinking the Inception Architecture for Computer Vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2017
Cited alongside, same era.
Learning to optimize tensor programs
Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Fast Adjustable Threshold For Uniform Neural Network Quantization (Winning solution of LPIRC-II)
Alexander Goncharenko, Andrey Denisov, Sergey Alyamkin, and Evgeny Terentev · 2018
Cited alongside, same era.
MobileNetV2: Inverted Residuals and Linear Bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang Chieh Chen · 2018
Later among the works it cites.
Low-bit Quantization of Neural Networks for Efficient Inference
Yoni Choukroun, Eli Kravchik, Fan Yang, and Pavel Kisilev · 2019
Later among the works it cites.
Same, Same But Different - Recovering Neural Network Quantization Error Through Weight Factorization
Eldad Meller, Alexander Finkelstein, Uri Almog, and Mark Grobman · 2019
Later among the works it cites.
EfficientNet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V. Le · 2019
Later among the works it cites.
Efficient Execution of Quantized Deep Learning Models: A Compiler Approach
Animesh Jain, Shoubhik Bhattacharya, Masahiro Masuda, Vin Sharma, and Yida Wang · 2020
Later among the works it cites.
Mlir: A compiler infrastructure for the end of moore’s law, 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko · 2018
Cited alongside, same era.
Glow: Graph lowering compiler techniques for neural networks
Nadav Rotem, Jordan Fix, Saleem Abdulrasool, Summer Deng, Roman Dzhabarov, James Hegeman, Roman Levenstein, Bert Maher, Nadathur Satish, Jakob Olesen, Jongsoo Park, Artem Rakhov, and Misha Smelyanskiy · 2018
Cited alongside, same era.
{ \{ TVM
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, et al
Cited in the paper.
Apache mxnet v1.2.0 optimized with intel® math kernel library for deep neural networks (intel® mkl-dnn)
Intel
Cited in the paper.
8 bit inference with TensorRT
Nvidia
Cited in the paper.
Chris Lattner, Mehdi Amini, Uday Bondhugula, Albert Cohen, Andy Davis, Jacques Pienaar, River Riddle, Tatiana Shpeisman, Nicolas Vasilache, and Oleksandr Zinenko · 2020
Later among the works it cites.
Nvidia cuda compiler
Wikipedia contributors · 2020
Later among the works it cites.