Fetching the paper…
Reading the bibliography…
Large Language Model training with 8-bit floating point (FP8) formats promises significant efficiency improvements, but reduced numerical precision makes training challenging.
Statistical Inference
Casella, G. and Berger, R. L · 2002
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Kingma, D. P. and Ba, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Mixed precision training
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., and Wu, H · 2018
Earlier work this paper cites.
Hybrid 8-bit floating point (hfp8) training and inference for deep neural networks
Sun, X., Choi, J., Chen, C.-Y., Wang, N., Venkataramani, S., Srinivasan, V. V., Cui, X., Zhang, W., and Gopalakrishnan, K · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
Tillet, P., Kung, H. T., and Cox, D · 2019
Earlier work this paper cites.
Fbgemm: Enabling high-performance low-precision deep learning inference, 2021
Khudia, D., Huang, J., Basu, P., Deng, S., Liu, H., Park, J., and Smelyanskiy, M · 2021
Earlier work this paper cites.
Composer
MosaicML · 2021
Earlier work this paper cites.
Tuning large neural networks via zero-shot hyperparameter transfer
Yang, G., Hu, E. J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J · 2021
Earlier work this paper cites.
LLM.int8(): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Swin transformer v2: Scaling up capacity and resolution
Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., Dong, L., et al · 2022
Cited alongside, same era.
Micikevicius, P., Stosic, D., Burgess, N., Cornea, M., Dubey, P., Grisenthwaite, R., Ha, S., Heinecke, A., Judd, P., Kamalu, J., et al · 2022
Cited alongside, same era.
Unit scaling: Out-of-the-box low-precision training
Blake, C., Orr, D., and Luschi, C · 2023
Cited alongside, same era.
Symbolic discovery of optimization algorithms
Chen, X., Liang, C., Huang, D., Real, E., Wang, K., Pham, H., Dong, X., Luong, T., Hsieh, C.-J., Lu, Y., and Le, Q. V · 2023
Cited alongside, same era.
Blazingly fast LLM evaluation for in-context learning, 2 2023
Dohmann, J · 2023
Cited alongside, same era.
TransformerEngine, 2023
u- μ \mu p: The unit-scaled maximal update parametrization
Blake, C., Eichenberg, C., Dean, J., Balles, L., Prince, L. Y., Deiseroth, B., Cruz-Salinas, A. F., Luschi, C., Weinbach, S., and Orr, D · 2024
Later among the works it cites.
Flex attention: A programming model for generating optimized attention kernels, 2024
Dong, J., Feng, B., Guessous, D., Liang, Y., and He, H · 2024
Later among the works it cites.
A large-scale exploration of μ \mu -transfer, 2024
Lingle, L · 2024
Later among the works it cites.
ReLU strikes back: Exploiting activation sparsity in large language models
Mirzadeh, S. I., Alizadeh-Vahid, K., Mehta, S., del Mundo, C. C., Tuzel, O., Samei, G., Rastegari, M., and Farajtabar, M · 2024
Later among the works it cites.
cuBLAS: cublasLtMatmul()
NVIDIA Corporation · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
NVIDIA · 2023
Cited alongside, same era.
A spectral condition for feature learning
Yang, G., Simon, J. B., and Bernstein, J · 2023
Cited alongside, same era.
Scaling FP8 training to trillion-token LLMs
Anonymous · 2024
Cited alongside, same era.
Calibrating the Mosaic evaluation Gauntlet, 4 2024
Barton, T · 2024
Cited alongside, same era.
LLM Foundry
MosaicML
Cited in the paper.
Streaming
MosaicML
Cited in the paper.
Asynchronous multiply-and-accumulate instruction: wgmma.mma_async
NVIDIA
Cited in the paper.
OLMo, T., Walsh, P., Soldaini, L., Groeneveld, D., Lo, K., Arora, S., Bhagia, A., Gu, Y., Huang, S., Jordan, M., et al · 2024
Later among the works it cites.
How to set AdamW’s weight decay as you scale model and dataset size, 2024
Wang, X. and Aitchison, L · 2024
Later among the works it cites.
Small-scale proxies for large-scale transformer training instabilities
Wortsman, M., Liu, P. J., Xiao, L., Everett, K. E., Alemi, A. A., Adlam, B., Co-Reyes, J. D., Gur, I., Kumar, A., Novak, R., Pennington, J., Sohl-Dickstein, J., Xu, K., Lee, J., Gilmer, J., and Kornblith, S · 2024
Later among the works it cites.
Tensor programs VI: Feature learning in infinite depth neural networks
Yang, G., Yu, D., Zhu, C., and Hayou, S · 2024
Later among the works it cites.