Fetching the paper…
Reading the bibliography…
Large language models (LLMs) excel in various tasks but face deployment challenges due to hardware constraints.
Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
Wu, H.; Judd, P.; Zhang, X.; Isaev, M.; and Micikevicius, P. 2020 · 2004
Earlier work this paper cites.
Pointer Sentinel Mixture Models
Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2007 · 2007
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L. u.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
SignSGD: Compressed Optimisation for Non-Convex Problems
Bernstein, J.; Wang, Y.-X.; Azizzadenesheli, K.; and Anandkumar, A. 2018 · 2018
Earlier work this paper cites.
Post Training 4-bit Quantization of Convolutional Networks for Rapid-Deployment
Banner, R.; Nahshan, Y.; and Soudry, D. 2019 · 2019
Earlier work this paper cites.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Gao, L.; Biderman, S.; Black, S.; Golding, L.; Hoppe, T.; Foster, C.; Phang, J.; He, H.; Thite, A.; Nabeshima, N.; et al. 2020 · 2020
Earlier work this paper cites.
Accurate Post Training Quantization with Small Calibration Sets
Hubara, I.; Nahshan, Y.; Hanani, Y.; Banner, R.; and Soudry, D. 2021 · 2021
Earlier work this paper cites.
Loss Aware Post-Training Quantization
Nahshan, Y.; Chmiel, B.; Baskin, C.; Zheltonozhskii, E.; Banner, R.; Bronstein, A. M.; and Mendelson, A. 2021 · 2021
Earlier work this paper cites.
FP8 Quantization: The Power of the Exponent
Kuzmin, A.; van Baalen, M.; Ren, Y.; Nagel, M.; Peters, J.; and Blankevoort, T. 2022 · 2022
Earlier work this paper cites.
Outlier Suppression: Pushing the Limit of Low-Bit Transformer Language Models
Wei, X.; Zhang, Y.; Zhang, X.; Gong, R.; Zhang, S.; Zhang, Q.; Yu, F.; and Liu, X. 2022 · 2022
Cited alongside, same era.
OPT: Open Pre-trained Transformer Language Models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; et al. 2022 · 2022
Cited alongside, same era.
QLoRA: Efficient Finetuning of Quantized LLMs
Dettmers, T.; Pagnoni, A.; Holtzman, A.; and Zettlemoyer, L. 2023 · 2023
Cited alongside, same era.
The Case for 4-Bit Precision: K-Bit Inference Scaling Laws
Dettmers, T.; and Zettlemoyer, L. 2023 · 2023
Cited alongside, same era.
SparseGPT: Massive Language Models Can be Accurately Pruned in One-Shot
Frantar, E.; and Alistarh, D. 2023 · 2023
Cited alongside, same era.
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
Xiao, G.; Lin, J.; Seznec, M.; Wu, H.; Demouth, J.; and Han, S. 2023 · 2023
Later among the works it cites.
NF4 Isn’t Information Theoretically Optimal (and That’s Good)
Yoshida, D. 2023 · 2023
Later among the works it cites.
QuantTune: Optimizing Model Quantization with Adaptive Outlier-Driven Fine Tuning
Chen, J.-M.; Chao, Y.-H.; Wang, Y.-J.; Shieh, M.-D.; Hsu, C.-C.; and Lin, W.-F. 2024 · 2024
Closest in time.
SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Dettmers, T.; Svirschevski, R.; Egiazarian, V.; Kuznedelev, D.; Frantar, E.; Ashkboos, S.; Borzunov, A.; Hoefler, T.; and Alistarh, D. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Frantar, E.; Ashkboos, S.; Hoefler, T.; and Alistarh, D. 2023 · 2023
Cited alongside, same era.
MiniLLM: Knowledge Distillation of Large Language Models
Gu, Y.; Dong, L.; Wei, F.; and Huang, M. 2023 · 2023
Cited alongside, same era.
OpenAI. 2023 · 2023
Cited alongside, same era.
Toward Accurate Post-Training Quantization for Image Super Resolution
Tu, Z.; Hu, J.; Chen, H.; and Wang, Y. 2023 · 2023
Cited alongside, same era.
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
Liu, S.-y.; Liu, Z.; Huang, X.; Dong, P.; and Cheng, K.-T. 2023a
Cited in the paper.
LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Liu, Z.; Oguz, B.; Zhao, C.; Chang, E.; Stock, P.; Mehdad, Y.; Shi, Y.; Krishnamoorthi, R.; and Chandra, V. 2023b
Cited in the paper.
LLaMA: Open and Efficient Foundation Language Models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023a
Cited in the paper.
Kim, S.; Hooper, C.; Gholami, A.; Dong, Z.; Li, X.; Shen, S.; Mahoney, M. W.; and Keutzer, K. 2024 · 2024
Closest in time.
OWQ: Outlier-Aware Weight Quantization for Efficient Fine-tuning and Inference of Large Language Models
Lee, C.; Jin, J.; Kim, T.; Kim, H.; and Park, E. 2024 · 2024
Closest in time.
AWQ: Activation-Aware Weight Quantization for LLM Compression and Acceleration
Lin, J.; Tang, J.; Tang, H.; Yang, S.; Dang, X.; and Han, S. 2024 · 2024
Closest in time.
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
Shao, W.; Chen, M.; Zhang, Z.; Xu, P.; Zhao, L.; Li, Z.; Zhang, K.; Gao, P.; Qiao, Y.; and Luo, P. 2024 · 2024
Closest in time.