Fetching the paper…
Reading the bibliography…
Neural network quantization has an inherent problem called accumulated quantization error, which is the key obstacle towards ultra-low precision, e.g., 2- or 3-bit precision.
Learning to forget: Continual prediction with LSTM
Felix A. Gers, Jürgen Schmidhuber, and Fred A. Cummins · 2000
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Çaglar Gülçehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Earlier work this paper cites.
Song Han, Huizi Mao, and William J. Dally · 2015
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara et al · 2016
Earlier work this paper cites.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher · 2016
Earlier work this paper cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2016
Earlier work this paper cites.
Xnor-net: Imagenet classification using binary convolutional neural networks
Mohammad Rastegari et al · 2016
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou et al · 2016
Earlier work this paper cites.
Chenzhuo Zhu et al · 2016
Cited alongside, same era.
Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks
Yu-Hsin Chen et al · 2017
Cited alongside, same era.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger · 2017
Cited alongside, same era.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob et al · 2017
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li et al · 2017
Cited alongside, same era.
High performance ultra-low-precision convolutions on mobile devices
Andrew Tulloch and Yangqing Jia · 2017
Later among the works it cites.
Balanced quantization: An effective and efficient approach to quantized neural networks
Shuchang Zhou et al · 2017
Later among the works it cites.
Arm compute library
ACL · 2018
Closest in time.
gemmlowp: a small self-contained low-precision gemm library
Benoit Jacob et al · 2018
Closest in time.
Zena: Zero-aware neural network accelerator
Dongyoung Kim, Junwhan Ahn, and Sungjoo Yoo · 2018
Closest in time.
Bi-real net: Enhancing the performance of 1-bit cnns with improved representational capability and advanced training algorithm
Zechun Liu et al · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
NVIDIA 8-bit inference width TensorRT
Szymon Migacz · 2017
Cited alongside, same era.
Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy
Asit K. Mishra and Debbie Marr · 2017
Cited alongside, same era.
WRPN: wide reduced-precision networks
Asit K. Mishra, Eriko Nurvitadhi, Jeffrey J. Cook, and Debbie Marr · 2017
Cited alongside, same era.
Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural networks
Hardik Sharma et al · 2017
Cited alongside, same era.
Bridging the accuracy gap for 2-bit quantized neural networks
Jungwook Choi et al
Cited in the paper.
Pact: Parameterized clipping activation for quantized neural networks
Jungwook Choi et al
Cited in the paper.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun
Cited in the paper.
Closest in time.
Energy-efficient neural network accelerator based on outlier-aware low-precision computation
Eunhyeok Park, Dongyoung Kim, and Sungjoo Yoo · 2018
Closest in time.
Towards effective low-bitwidth convolutional neural networks
Bohan Zhuang et al · 2018
Closest in time.