Fetching the paper…
Reading the bibliography…
We propose a novel value-aware quantization which applies aggressively reduced precision to the majority of data while separately handling a small amount of large data in high precision, which reduces total quantization errors under very low precision.
Marcus, M.P., Marcinkiewicz, M.A., Santorini, B.: Building a large annotated corpus of English: The Penn Treebank. Computational linguistics 19(2), 313–330 (1993)
1993
Earlier work this paper cites.
Hong, S., Kim, H.: An integrated GPU power and performance model. In: International Symposium on Computer Architecture (ISCA). pp. 280–289 (2010)
2010
Earlier work this paper cites.
Marcel, S., Rodriguez, Y.: Torchvision the machine-vision package of torch. ACM Multimedia (2010)
2010
Earlier work this paper cites.
He, X., et al.: Practical lessons from predicting clicks on ads at facebook. In: International Workshop on Data Mining for Online Advertising (ADKDD). pp. 1–9. ACM (2014)
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Chen, T., et al.: Training deep nets with sublinear memory cost. arXiv:1604.06174 (2016)
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Rastegari, M., et al.: Xnor-net: Imagenet classification using binary convolutional neural networks. the European Conference on Computer Vision (ECCV) pp. 525–542 (2016)
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Zhu, C., et al.: Trained ternary quantization. arXiv:1612.01064 (2016)
2016
Cited alongside, same era.
Bakunas-Milanowski, D., et al.: Efficient algorithms for stream compaction on gpus. International Journal of Networking and Computing (IJNC) 7(2), 208–226 (2017)
2017
Cited alongside, same era.
Ginsburg, B., et al.: NVIDIA Mixed Precision Training on Volta GPUs. GPU Technology Conference (2017)
Park, E., Ahn, J., Yoo, S.: Weighted-entropy-based quantization for deep neural networks. Computer Vision and Pattern Recognition (CVPR) pp. 7197–7205 (2017)
2017
Later among the works it cites.
Paszke, A., et al.: Pytorch (2017)
2017
Later among the works it cites.
Press, O., Wolf, L.: Using the output embedding to improve language models. In: the European Chapter of the Association for Computational Linguistics (EACL). pp. 157–163 (2017)
2017
Later among the works it cites.
Umuroglu, Y., et al.: Finn: A framework for fast, scalable binarized neural network inference. In: ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA). pp. 65–74 (2017)
2017
Later among the works it cites.
Zhou, S., et al.: Balanced quantization: An effective and efficient approach to quantized neural networks. J. Comput. Sci. Technol. 32(4), 667–682 (2017)
2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Gomez, A.N., et al.: The reversible residual network: Backpropagation without storing activations. In: Advances in Neural Information Processing Systems (NIPS). pp. 2211–2221 (2017)
2017
Cited alongside, same era.
Jouppi, N.P., et al.: In-datacenter performance analysis of a tensor processing unit. In: International Symposium on Computer Architecture (ISCA). pp. 1–12 (2017)
2017
Cited alongside, same era.
Migacz, S.: NVIDIA 8-bit inference width TensorRT. GPU Technology Conference (2017)
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Mishra, A., et al.: Wrpn: Wide reduced-precision networks. arXiv:1709.01134 (2017)
2017
Cited alongside, same era.
Later among the works it cites.
De Sa, C., et al.: High-accuracy low-precision training. arXiv:1803.03383 (2018)
2018
Closest in time.
Hazelwood, K., et al.: Applied machine learning at facebook: A datacenter infrastructure perspective. International Symposium on High-Performance Computer Architecture (HPCA) (2018)
2018
Closest in time.
Jia, Y., Peter, V.: Delivering real-time ai in the palm of your hand. https://code.facebook.com/posts/196146247499076/delivering-real-time-ai-in-the-palm-of-your-hand/
2018
Closest in time.
Polino, A., Pascanu, R., Alistarh, D.: Model compression via distillation and quantization. International Conference on Learning Representation (ICLR) (2018)
2018
Closest in time.