Fetching the paper…
Reading the bibliography…
Low-precision training is considered an effective strategy for reducing both training and downstream inference costs.
Ieee standard 754 for binary floating-point arithmetic
Kahan, W · 1996
Earlier work this paper cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., et al · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Hybrid 8-bit floating point (hfp8) training and inference for deep neural networks
Sun, X., Choi, J., Chen, C.-Y., Wang, N., Venkataramani, S., Srinivasan, V. V., Cui, X., Zhang, W., and Gopalakrishnan, K · 2019
Earlier work this paper cites.
Qpytorch: A low-precision arithmetic simulation framework
Zhang, T., Lin, Z., Yang, G., and De Sa, C · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Earlier work this paper cites.
Fp8 quantization: The power of the exponent
Kuzmin, A., van Baalen, M., Ren, Y., Nagel, M., Peters, J., and Blankevoort, T · 2022
Earlier work this paper cites.
Micikevicius, P., Stosic, D., Burgess, N., Cornea, M., Dubey, P., Grisenthwaite, R., Ha, S., Heinecke, A., Judd, P., Kamalu, J., et al · 2022
Earlier work this paper cites.
Quantease: Optimization-based quantization for language models-an efficient and intuitive algorithm
Behdin, K., Acharya, A., Aman Gupta, S. K., and Mazumder, R · 2023
Earlier work this paper cites.
The case for 4-bit precision: k-bit inference scaling laws
Dettmers, T. and Zettlemoyer, L · 2023
Earlier work this paper cites.
Fp8-lm: Training fp8 large language models
Peng, H., Wu, K., Wei, Y., Zhao, G., Yang, Y., Liu, Z., Xiong, Y., Yang, Z., Ni, B., Hu, J., et al · 2023
Cited alongside, same era.
Bitnet: Scaling 1-bit transformers for large language models
Wang, H., Ma, S., Dong, L., Huang, S., Wang, H., Ma, L., Yang, F., Wang, R., Wu, Y., and Wei, F · 2023
Cited alongside, same era.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Cited alongside, same era.
Nf4 isn’t information theoretically optimal (and that’s good)
Yoshida, D · 2023
Cited alongside, same era.
Revisiting block-based quantisation: What is important for sub-8-bit llm inference?
Zhang, C., Cheng, J., Shumailov, I., Constantinides, G. A., and Zhao, Y · 2023
Kumar, T., Ankner, Z., Spector, B. F., Bordelon, B., Muennighoff, N., Paul, M., Pehlevan, C., Ré, C., and Raghunathan, A · 2024
Later among the works it cites.
A comprehensive study on quantization techniques for large language models
Lang, J., Guo, Z., and Huang, S · 2024
Later among the works it cites.
Surge phenomenon in optimal learning rate and batch size scaling
Li, S., Zhao, P., Zhang, H., Sun, S., Wu, H., Jiao, D., Wang, W., Liu, C., Fang, Z., Xue, J., Tao, Y., CUI, B., and Wang, D · 2024
Later among the works it cites.
Scaling laws in linear regression: Compute, parameters, and data
Lin, L., Wu, J., Kakade, S. M., Bartlett, P. L., and Lee, J. D · 2024
Later among the works it cites.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model
Liu, A., Feng, B., Wang, B., Wang, B., Liu, B., Zhao, C., Dengr, C., Ruan, C., Dai, D., Guo, D., et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Explaining neural scaling laws
Bahri, Y., Dyer, E., Kaplan, J., Lee, J., and Sharma, U · 2024
Cited alongside, same era.
Transformers are ssms: Generalized models and efficient algorithms through structured state space duality
Dao, T. and Gu, A · 2024
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Extreme compression of large language models via additive quantization
Egiazarian, V., Panferov, A., Kuznedelev, D., Frantar, E., Babenko, A., and Alistarh, D · 2024
Cited alongside, same era.
Olmo: Accelerating the science of language models
Groeneveld, D., Beltagy, I., Walsh, P., Bhagia, A., Kinney, R., Tafjord, O., Jha, A. H., Ivison, H., Magnusson, I., Wang, Y., Arora, S., Atkinson, D., Authur, R., Chandu, K., Cohan, A., Dumas, J., Elazar, Y., Gu, Y., Hessel, J., Khot, T., Merrill, W., Morrison, J., Muennighoff, N., Naik, A., Nam, C., Peters, M. E., Pyatkin, V., Ravichander, A., Schwenk, D., Shah, S., Smith, W., Subramani, N., Wortsman, M., Dasigi, P., Lambert, N., Richardson, K., Dodge, J., Lo, K., Soldaini, L., Smith, N. A., and Hajishirzi, H · 2024
Cited alongside, same era.
Later among the works it cites.
The era of 1-bit llms: All large language models are in 1.58 bits
Ma, S., Wang, H., Ma, L., Wang, L., Wang, W., Huang, S., Dong, L., Wang, R., Xue, J., and Wei, F · 2024
Later among the works it cites.
Ouyang, X., Ge, T., Hartvigsen, T., Zhang, Z., Mi, H., and Yu, D · 2024
Later among the works it cites.
Exploring quantization techniques for large-scale language models: Methods, challenges and future directions
Shen, A., Lai, Z., and Li, D · 2024
Later among the works it cites.
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Soldaini, L., Kinney, R., Bhagia, A., Schwenk, D., Atkinson, D., Authur, R., Bogin, B., Chandu, K., Dumas, J., Elazar, Y., Hofmann, V., Jha, A. H., Kumar, S., Lucy, L., Lyu, X., Lambert, N., Magnusson, I., Morrison, J., Muennighoff, N., Naik, A., Nam, C., Peters, M. E., Ravichander, A., Richardson, K., Shen, Z., Strubell, E., Subramani, N., Tafjord, O., Walsh, P., Zettlemoyer, L., Smith, N. A., Hajishirzi, H., Beltagy, I., Groeneveld, D., Dodge, J., and Lo, K · 2024
Later among the works it cites.
Hunyuan-large: An open-source moe model with 52 billion activated parameters by tencent
Sun, X., Chen, Y., Huang, Y., Xie, R., Zhu, J., Zhang, K., Li, S., Yang, Z., Han, J., Shu, X., et al · 2024
Later among the works it cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al · 2024
Later among the works it cites.