Fetching the paper…
Reading the bibliography…
Pretraining transformers are generally time-consuming.
Hawq-v2: Hessian aware trace-weighted quantization of neural networks
Dong, Z., Yao, Z., Cai, Y., Arfeen, D., Gholami, A., Mahoney, M. W., and Keutzer, K · 1911
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
Bojar, O., Buck, C., Federmann, C., Haddow, B., Koehn, P., Leveling, J., Monz, C., Pecina, P., Post, M., Saint-Amand, H., et al · 2014
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Scalable methods for 8-bit training of neural networks
Banner, R., Hubara, I., Hoffer, E., and Soudry, D · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., and Kalenichenko, D · 2018
Earlier work this paper cites.
Nvidia tensor core programmability, performance & precision
Markidis, S., Der Chien, S. W., Laure, E., Peng, I. B., and Vetter, J. S · 2018
Earlier work this paper cites.
Mixed precision training
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., et al · 2018
Earlier work this paper cites.
Learned step size quantization
Esser, S. K., McKinstry, J. L., Bablani, D., Appuswamy, R., and Modha, D. S · 2019
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout
Fan, A., Grave, E., and Joulin, A · 2019
Earlier work this paper cites.
Open clone of openai’s unreleased webtext dataset scraper, 2019
Peterson, J., Meylan, S., and Bourgin, D · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Q-bert: Hessian based ultra low precision quantization of bert
Shen, S., Dong, Z., Ye, J., Ma, L., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2019
Cited alongside, same era.
Hybrid 8-bit floating point (hfp8) training and inference for deep neural networks
Sun, X., Choi, J., Chen, C.-Y., Wang, N., Venkataramani, S., Srinivasan, V. V., Cui, X., Zhang, W., and Gopalakrishnan, K · 2019
Cited alongside, same era.
Triton: an intermediate language and compiler for tiled neural network computations
Tillet, P., Kung, H.-T., and Cox, D · 2019
Cited alongside, same era.
Binarybert: Pushing the limit of bert quantization
Bai, H., Zhang, W., Hou, L., Shang, L., Jin, J., Jiang, X., Liu, Q., Lyu, M., and King, I · 2020
Cited alongside, same era.
A statistical framework for low-bitwidth training of deep neural networks
Chen, J., Gai, Y., Yao, Z., Mahoney, M. W., and Gonzalez, J. E · 2020
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Later among the works it cites.
Llm. int8 (): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Later among the works it cites.
Nvidia transformer engine
Nvidia · 2022
Later among the works it cites.
Mkq-bert: Quantized bert with 4-bits weights and activations
Tang, H., Zhang, X., Liu, K., Zhu, J., and Kang, Z · 2022
Later among the works it cites.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ultra-low precision 4-bit training of deep neural networks
Sun, X., Wang, N., Chen, C.-Y., Ni, J., Agrawal, A., Cui, X., Venkataramani, S., El Maghraoui, K., Srinivasan, V. V., and Gopalakrishnan, K · 2020
Cited alongside, same era.
Ternarybert: Distillation-aware ultra-low bit bert
Zhang, W., Hou, L., Yin, Y., Shang, L., Chen, X., Jiang, X., and Liu, Q · 2020
Cited alongside, same era.
Towards unified int8 training for convolutional neural network
Zhu, F., Gong, R., Yu, F., Liu, X., Wang, Y., Li, Z., Yang, X., and Yan, J · 2020
Cited alongside, same era.
Logarithmic unbiased quantization: Practical 4-bit training in deep learning
Chmiel, B., Banner, R., Hoffer, E., Yaacov, H. B., and Soudry, D · 2021
Cited alongside, same era.
I-bert: Integer-only bert quantization
Kim, S., Gholami, A., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Cited alongside, same era.
Post-training quantization for vision transformer
Liu, Z., Wang, Y., Han, K., Zhang, W., Ma, S., and Gao, W · 2021
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2021
Cited alongside, same era.
Quip: 2-bit quantization of large language models with guarantees
Chee, J., Cai, Y., Kuleshov, V., and De Sa, C · 2023
Later among the works it cites.
Squeezellm: Dense-and-sparse quantization
Kim, S., Hooper, C., Gholami, A., Dong, Z., Li, X., Shen, S., Mahoney, M. W., and Keutzer, K · 2023
Later among the works it cites.
Awq: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Fp8-lm: Training fp8 large language models
Peng, H., Wu, K., Wei, Y., Zhao, G., Yang, Y., Liu, Z., Xiong, Y., Yang, Z., Ni, B., Hu, J., et al · 2023
Later among the works it cites.
Training and inference of large language models using 8-bit floating point
Perez, S. P., Zhang, Y., Briggs, J., Blake, C., Levy-Kramer, J., Balanca, P., Luschi, C., Barlow, S., and Fitzgibbon, A. W · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Stable and low-precision training for large-scale vision-language models
Wortsman, M., Dettmers, T., Zettlemoyer, L., Morcos, A., Farhadi, A., and Schmidt, L · 2023
Later among the works it cites.
Training transformers with 4-bit integers
Xi, H., Li, C., Chen, J., and Zhu, J · 2023
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Later among the works it cites.