Fetching the paper…
Reading the bibliography…
State-of-the-art generic low-precision training algorithms use a mix of 16-bit and 32-bit precision, creating the folklore that 16-bit hardware compute units alone are not enough to maximize model accuracy.
Deep learning recommendation model for personalization and recommendation systems
Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G. Azzolini, Dmytro Dzhulgakov, Andrey Mallevich, Ilia Cherniavskii, Yinghai Lu, Raghuraman Krishnamoorthi, Ansha Yu, Volodymyr Kondratenko, Stephanie Pereira, Xianjie Chen, Wenlin Chen, Vijay Rao, Bill Jia, Liang Xiong, and Misha Smelyanskiy · 1906
Earlier work this paper cites.
Further remarks on reducing truncation errors, commun
W Kahan · 1965
Earlier work this paper cites.
The cifar-10 dataset, 2009
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Fpu generator for design space exploration
Sameh Galal, Ofer Shacham, John S Brunhaver II, Jing Pu, Artem Vassiliev, and Mark Horowitz · 2013
Earlier work this paper cites.
Introduction to numerical analysis , volume 12
Josef Stoer and Roland Bulirsch · 2013
Earlier work this paper cites.
1.1 computing’s energy problem (and what we can do about it)
Mark Horowitz · 2014
Earlier work this paper cites.
Display advertising challenge, 2014
CriteoLabs · 2014
Earlier work this paper cites.
Training deep neural networks with low precision multiplications
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2014
Earlier work this paper cites.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Earlier work this paper cites.
Taming the wild: A unified analysis of hogwild-style algorithms
Christopher M De Sa, Ce Zhang, Kunle Olukotun, and Christopher Ré · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou · 2016
Cited alongside, same era.
Recurrent neural networks with limited numerical precision
Joachim Ott, Zhouhan Lin, Ying Zhang, Shih-Chii Liu, and Yoshua Bengio · 2016
Cited alongside, same era.
Understanding and optimizing asynchronous low-precision stochastic gradient descent
Christopher De Sa, Matthew Feldman, Christopher Ré, and Kunle Olukotun · 2017
Cited alongside, same era.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2017
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Training with low-precision embedding tables
Jian Zhang, Jiyan Yang, and Hector Yuen · 2018
Later among the works it cites.
Criteo 1tb click logs dataset, 2018
CriteoLabs · 2018
Later among the works it cites.
Megatron-lm: Training multi-billion parameter language models using gpu model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Later among the works it cites.
Zero: Memory optimization towards training a trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al · 2017
Cited alongside, same era.
Zipml: Training linear models with end-to-end low precision, and a little bit of deep learning
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit
Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, et al · 2017
Cited alongside, same era.
Training quantized nets: A deeper understanding
Hao Li, Soham De, Zheng Xu, Christoph Studer, Hanan Samet, and Tom Goldstein · 2017
Cited alongside, same era.
Reducing and monitoring round-off error propagation for symplectic implicit runge-kutta schemes
Mikel Antonana, Joseba Makazaga, and Ander Murua · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Bfloat16-hardware numerics definition, 2018
Intel · 2018
Cited alongside, same era.
Analysis of quantized models
Lu Hou, Ruiliang Zhang, and James T Kwok · 2018
Cited alongside, same era.
Regularized evolution for image classifier architecture search
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le · 2019
Later among the works it cites.
A study of bfloat16 for deep learning training
Dhiraj Kalamkar, Dheevatsa Mudigere, Naveen Mellempudi, Dipankar Das, Kunal Banerjee, Sasikanth Avancha, Dharma Teja Vooturi, Nataraj Jammalamadaka, Jianyu Huang, Hector Yuen, et al · 2019
Later among the works it cites.
Bfloat16 processing for neural networks
Neil Burgess, Jelena Milanovic, Nigel Stephens, Konstantinos Monachopoulos, and David Mansell · 2019
Later among the works it cites.
Cloud tpu: Codesigning architecture and infrastructure
C Chao and B Saeta · 2019
Later among the works it cites.
Bfloat16 processing for neural networks on armv8-a
N Stephens · 2019
Later among the works it cites.
QPyTorch: A low-precision arithmetic simulation framework
Tianyi Zhang, Zhiqiu Lin, Guandao Yang, and Christopher De Sa · 2019
Later among the works it cites.
Hybrid 8-bit floating point (hfp8) training and inference for deep neural networks
Xiao Sun, Jungwook Choi, Chia-Yu Chen, Naigang Wang, Swagath Venkataramani, Vijayalakshmi Viji Srinivasan, Xiaodong Cui, Wei Zhang, and Kailash Gopalakrishnan · 2019
Later among the works it cites.
Floating point and ieee 754 compliance for nvidia gpus, 2020
Nvidia · 2020
Closest in time.
Stochastic rounding and reduced-precision fixed-point arithmetic for solving neural ordinary differential equations
Michael Hopkins, Mantas Mikaitis, Dave R Lester, and Steve Furber · 2020
Closest in time.