Fetching the paper…
Reading the bibliography…
Deep learning models typically use single-precision (FP32) floating point data types for representing activations and weights, but a slew of recent research work has shown that computations with reduced-precision data types (FP16, 16-bit integers, 8-bit integers or even 4- or 2-bit integers) are enough to achieve same accuracy as FP32 and are much more efficient.
Anatomy of High-performance Matrix Multiplication
Kazushige Goto and Robert A. van de Geijn. 2008 · 2008
Earlier work this paper cites.
Roofline: An Insightful Visual Performance Model for Multicore Architectures
Samuel Williams, Andrew Waterman, and David Patterson. 2009 · 2009
Earlier work this paper cites.
BLIS: A Framework for Rapidly Instantiating BLAS Functionality
Field G. Van Zee and Robert A. van de Geijn. 2015 · 2015
Earlier work this paper cites.
The BLIS Framework: Experiments in Portability
Field G. Van Zee, Tyler M. Smith, Bryan Marker, Tze Meng Low, Robert A. Van De Geijn, Francisco D. Igual, Mikhail Smelyanskiy, Xianyi Zhang, Michael Kistler, Vernon Austel, John A. Gunnels, and Lee Killough. 2016 · 2016
Cited alongside, same era.
gemmlowp: a small self-contained low-precision GEMM library
Google. 2017 · 2017
Cited alongside, same era.
Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew G. Howard, Hartwig Adam, and Dmitry Kalenichenko. 2017 · 2017
Cited alongside, same era.
PyTorch 1.0
Facebook Open Source. 2018b
Cited in the paper.
Rosetta: Understanding text in images and videos with machine learning
Facebook Engineering Blog. 2018 · 2018
Later among the works it cites.
Intel Xeon Processor E5-2680 v4
Intel. 2018 · 2018
Later among the works it cites.
Energy-Efficient Neural Network Accelerator Based on Outlier-Aware Low-Precision Computation. In 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA) . 688–698
E. Park, D. Kim, and S. Yoo. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…