Fetching the paper…
Reading the bibliography…
Given the current trend of increasing size and complexity of machine learning architectures, it has become of critical importance to identify new approaches to improve the computational efficiency of model training.
On a test of whether one of two random variables is stochastically larger than the other
Henry B. Mann and Donald R. Whitney · 1947
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T. Poliak · 1964
Earlier work this paper cites.
Statistical theory of quantization
Bernard Widrow, István Kollár, and Ming-Chang Liu · 1996
Earlier work this paper cites.
Quantization Noise: Roundoff Error in Digital Computation, Signal Processing, Control, and Communications
Bernard Widrow and István Kollár · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Lecture 6.5-rmsprop, Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Computer arithmetic and validity: theory, implementation, and applications , volume 33
Ulrich Kulisch · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Findings of the 2014 Workshop on Statistical Machine Translation
Ondrej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, et al · 2014
Earlier work this paper cites.
Training deep neural networks with low precision multiplications
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
BinaryConnect: Training deep neural networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Yukun Zhu, Ryan Kiros, Richard S. Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Cited alongside, same era.
Findings of the 2016 Conference on Machine Translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, et al · 2016
Cited alongside, same era.
Matthieu Courbariaux, Itay Hubara, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Cited alongside, same era.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Quantized neural networks: Training neural networks with low precision weights and activations
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Later among the works it cites.
Rethinking floating point for deep learning
Jeff Johnson · 2018
Later among the works it cites.
Mixed-precision training for NLP and speech recognition with OpenSeq2Seq
Oleksii Kuchaiev, Boris Ginsburg, Igor Gitman, Vitaly Lavrukhin, Jason Li, Huyen Nguyen, Carl Case, and Paulius Micikevicius · 2018
Later among the works it cites.
A call for clarity in reporting BLEU scores
Matt Post · 2018
Later among the works it cites.
Training deep neural networks with 8-bit floating point numbers
Naigang Wang, Jungwook Choi, Daniel Brand, Chia-Yu Chen, and Kailash Gopalakrishnan · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Cited alongside, same era.
Convolutional neural networks using logarithmic data representation
Daisuke Miyashita, Edward H. Lee, and Boris Murmann · 2016
Cited alongside, same era.
DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou · 2016
Cited alongside, same era.
Beating floating point at its own game: Posit arithmetic
John L. Gustafson and Isaac Yonemoto · 2017
Cited alongside, same era.
Fixing weight decay regularization in Adam
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu · 2017
Cited alongside, same era.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Later among the works it cites.
DLFloat: A 16-bit floating point format designed for deep learning training and inference
Ankur Agrawal, Silvia M. Mueller, Bruce M. Fleischer, Jungwook Choi, Naigang Wang, Xiao Sun, and Kailash Gopalakrishnan · 2019
Later among the works it cites.
A study of BFLOAT16 for deep learning training
Dhiraj Kalamkar, Dheevatsa Mudigere, Naveen Mellempudi, Dipankar Das, Kunal Banerjee, Sasikanth Avancha, Dharma Teja Vooturi, Nataraj Jammalamadaka, Jianyu Huang, Hector Yuen, Jiyan Yang, Jongsoo Park, Alexander Heinecke, Evangelos Georganas, Sudarshan Srinivasan, Abhisek Kundu, Misha Smelyanskiy, Bharat Kaul, and Pradeep Dubey · 2019
Later among the works it cites.
Mixed precision training with 8-bit floating point
Naveen Mellempudi, Sudarshan Srinivasan, Dipankar Das, and Bharat Kaul · 2019
Later among the works it cites.
Fully quantized transformer for improved translation
Gabriele Prato, Ella Charlaix, and Mehdi Rezagholizadeh · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Later among the works it cites.
Hybrid 8-bit floating point (HFP8) training and inference for deep neural networks
Xiao Sun, Jungwook Choi, Chia-Yu Chen, Naigang Wang, Swagath Venkataramani, Vijayalakshmi Srinivasan, Xiaodong Cui, Wei Zhang, and Kailash Gopalakrishnan · 2019
Later among the works it cites.
EfficientNet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V. Le · 2019
Later among the works it cites.
CutMix: Regularization strategy to train strong classifiers with localizable features
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo · 2019
Later among the works it cites.
Adaptive loss scaling for mixed precision training
Ruizhe Zhao, Brian Vogel, and Tanvir Ahmed · 2019
Later among the works it cites.
Dominic Masters, Antoine Labatie, Zach Eaton-Rosen, and Carlo Luschi · 2021
Later among the works it cites.