Fetching the paper…
Reading the bibliography…
FP8 formats are gaining popularity to boost the computational efficiency for training and inference of large deep learning models.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, et al · 2017
Earlier work this paper cites.
Mixed-precision training for nlp and speech recognition with openseq2seq
O. Kuchaiev, B. Ginsburg, I. Gitman, V. Lavrukhin, J. Li, H. Nguyen, C. Case, and P. Micikevicius · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2018
Earlier work this paper cites.
Ieee standard for floating-point arithmetic, 2019
C. S. IEEE · 2019
Earlier work this paper cites.
Dissecting the graphcore ipu architecture via microbenchmarking
Z. Jia, B. Tillman, M. Maggioni, and D. P. Scarpazza · 2019
Earlier work this paper cites.
A study of bfloat16 for deep learning training
D. Kalamkar, D. Mudigere, N. Mellempudi, D. Das, K. Banerjee, S. Avancha, D. T. Vooturi, N. Jammalamadaka, J. Huang, H. Yuen, et al · 2019
Earlier work this paper cites.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Understanding and overcoming the challenges of efficient transformer quantization
Y. Bondarenko, M. Nagel, and T. Blankevoort · 2021
Earlier work this paper cites.
A framework for few-shot language model evaluation, Sept. 2021
L. Gao, J. Tow, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, K. McDonell, N. Muennighoff, J. Phang, L. Reynolds, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou · 2021
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training gopher
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, et al · 2021
Earlier work this paper cites.
Tuning large neural networks via zero-shot hyperparameter transfer
G. Yang, E. Hu, I. Babuschkin, S. Sidor, X. Liu, D. Farhi, N. Ryder, J. Pachocki, W. Chen, and J. Gao · 2021
Earlier work this paper cites.
The case for 4-bit precision: k-bit inference scaling laws
T. Dettmers and L. Zettlemoyer · 2022
Cited alongside, same era.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh · 2022
Cited alongside, same era.
Fp8 quantization: The power of the exponent
A. Kuzmin, M. Van Baalen, Y. Ren, M. Nagel, J. Peters, and T. Blankevoort · 2022
Cited alongside, same era.
Bow-2000 ipu-machine datasheet, 2022a
Graphcore · 2023
Closest in time.
Graphcore tile vertex isa release 1.3.1 ipu21, 2022b
Graphcore · 2023
Closest in time.
Nvidia h100 tensor core gpu architecture, 2022a
Nvidia · 2023
Closest in time.
Transformer engine
Nvidia · 2023
Closest in time.
Interim report on 8-bit binary floating-point formats
I. W. G. P3109 · 2023
Closest in time.
Training large models more stably with automatic loss scaling, 2022
S. P. Perez · 2023
Closest in time.
Efficiently scaling transformer inference
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, J. Heek, K. Xiao, S. Agrawal, and J. Dean · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Micikevicius, D. Stosic, N. Burgess, M. Cornea, P. Dubey, R. Grisenthwaite, S. Ha, A. Heinecke, P. Judd, J. Kamalu, et al · 2022
Cited alongside, same era.
8-bit numerical formats for deep neural networks
B. Noune, P. Jones, D. Justus, D. Masters, and C. Luschi · 2022
Cited alongside, same era.
Bloom: A 176b-parameter open-access multilingual language model
T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ilić, D. Hesslow, R. Castagné, A. S. Luccioni, F. Yvon, M. Gallé, et al · 2022
Cited alongside, same era.
Intriguing properties of quantization at scale
A. Ahmadian, S. Dash, H. Chen, B. Venkitesh, S. Gou, P. Blunsom, A. Üstün, and S. Hooker · 2023
Cited alongside, same era.
Unit scaling: Out-of-the-box low-precision training
C. Blake, D. Orr, and C. Luschi · 2023
Cited alongside, same era.
Quantizable transformers: Removing outliers by helping attention heads do nothing
Y. Bondarenko, M. Nagel, and T. Blankevoort · 2023
Cited alongside, same era.
Transformer inference arithmetic, 2022
C. Chen · 2023
Cited alongside, same era.
Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster
N. Dey, G. Gosal, H. Khachane, W. Marshall, R. Pathria, M. Tom, J. Hestness, et al · 2023
Cited alongside, same era.
Closest in time.
Why gpt-3.5 is (mostly) cheaper than llama 2, 2023
A. Sanger · 2023
Closest in time.
A guide to tesla’s configurable floating point formats & arithmetic
Tesla · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Closest in time.
Fp8 versus int8 for efficient deep learning inference, 2023
M. van Baalen, A. Kuzmin, S. S. Nair, Y. Ren, E. Mahurin, C. Patel, S. Subramanian, S. Lee, M. Nagel, J. Soriaga, and T. Blankevoort · 2023
Closest in time.
Smoothquant: Accurate and efficient post-training quantization for large language models
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han · 2023
Closest in time.