Fetching the paper…
Reading the bibliography…
Quantization stands as a pivotal technique for large language model (LLM) serving, yet it poses significant challenges particularly in achieving effective low-bit quantization.
Hellaswag: Can a machine really finish your sentence?
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 1905
Earlier work this paper cites.
Topics in Matrix Analysis
Horn, R. A.; and Johnson, C. R. 1991 · 1991
Earlier work this paper cites.
The penn treebank: Annotating predicate argument structure
Marcus, M.; Kim, G.; Marcinkiewicz, M. A.; MacIntyre, R.; Bies, A.; Ferguson, M.; Katz, K.; and Schasberger, B. 1994 · 1994
Earlier work this paper cites.
Matrix Analysis and Applied Linear Algebra Book and Solutions Manual
Meyer, C. D. 2000 · 2000
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S.; Li, B. Z.; Khabsa, M.; Fang, H.; and Ma, H. 2020 · 2006
Earlier work this paper cites.
The effective rank: A measure of effective dimensionality
Roy, O.; and Vetterli, M. 2007 · 2007
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020 · 2009
Earlier work this paper cites.
Contrastive distillation on intermediate representations for language model compression
Sun, S.; Gan, Z.; Cheng, Y.; Fang, Y.; Wang, S.; and Liu, J. 2020 · 2009
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y.; Zellers, R.; Gao, J.; Choi, Y.; et al. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
Evaluating Large Language Models Trained on Code
Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; de Oliveira Pinto, H. P.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; Ray, A.; Puri, R.; Krueger, G.; Petrov, M.; Khlaaf, H.; Sastry, G.; Mishkin, P.; Chan, B.; Gray, S.; Ryder, N.; Pavlov, M.; Power, A.; Kaiser, L.; Bavarian, M.; Winter, C.; Tillet, P.; Such, F. P.; Cummings, D.; Plappert, M.; Chantzis, F.; Barnes, E.; Herbert-Voss, A.; Guss, W. H.; Nichol, A.; Paino, A.; Tezak, N.; Tang, J.; Babuschkin, I.; Balaji, S.; Jain, S.; Saunders, W.; Hesse, C.; Carr, A. N.; Leike, J.; Achiam, J.; Misra, V.; Morikawa, E.; Radford, A.; Knight, M.; Brundage, M.; Murati, M.; Mayer, K.; Welinder, P.; McGrew, B.; Amodei, D.; McCandlish, S.; Sutskever, I.; and Zaremba, W. 2021 · 2021
Cited alongside, same era.
Training Verifiers to Solve Math Word Problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; Hesse, C.; and Schulman, J. 2021 · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Cited alongside, same era.
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G.; Lin, J.; Seznec, M.; Wu, H.; Demouth, J.; and Han, S. 2023 · 2023
Later among the works it cites.
Compressing transformers: features are low-rank, but weights are not!
Yu, H.; and Wu, J. 2023 · 2023
Later among the works it cites.
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
DeepSeek-AI. 2024 · 2024
Closest in time.
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
Lin, J.; Tang, J.; Tang, H.; Yang, S.; Chen, W.-M.; Wang, W.-C.; Xiao, G.; Dang, X.; Gan, C.; and Han, S. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dettmers, T.; Lewis, M.; Belkada, Y.; and Zettlemoyer, L. 2022 · 2022
Cited alongside, same era.
Optimal brain compression: A framework for accurate post-training quantization and pruning
Frantar, E.; and Alistarh, D. 2022 · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E.; Ashkboos, S.; Hoefler, T.; and Alistarh, D. 2022 · 2022
Cited alongside, same era.
The optimal bert surgeon: Scalable and accurate second-order pruning for large language models
Kurtic, E.; Campos, D.; Nguyen, T.; Frantar, E.; Kurtz, M.; Fineran, B.; Goin, M.; and Alistarh, D. 2022 · 2022
Cited alongside, same era.
Bai, J.; Bai, S.; Chu, Y.; Cui, Z.; Dang, K.; Deng, X.; Fan, Y.; Ge, W.; Han, Y.; Huang, F.; et al. 2023 · 2023
Cited alongside, same era.
Lq-lora: Low-rank plus quantized matrix decomposition for efficient language model finetuning
Guo, H.; Greengard, P.; Xing, E. P.; and Kim, Y. 2023 · 2023
Cited alongside, same era.
Loftq: Lora-fine-tuning-aware quantization for large language models
Li, Y.; Yu, Y.; Liang, C.; He, P.; Karampatziakis, N.; Chen, W.; and Zhao, T. 2023 · 2023
Cited alongside, same era.
Llm-pruner: On the structural pruning of large language models
Ma, X.; Fang, G.; and Wang, X. 2023 · 2023
Cited alongside, same era.
OpenCompass: A Universal Evaluation Platform for Foundation Models
OpenCompass. 2023 · 2023
Cited alongside, same era.
Closest in time.
SpinQuant–LLM quantization with learned rotations
Liu, Z.; Zhao, C.; Fedorov, I.; Soran, B.; Choudhary, D.; Krishnamoorthi, R.; Chandra, V.; Tian, Y.; and Blankevoort, T. 2024 · 2024
Closest in time.
Llama Team, A. . M. 2024 · 2024
Closest in time.
Are emergent abilities of large language models a mirage?
Schaeffer, R.; Miranda, B.; and Koyejo, S. 2024 · 2024
Closest in time.
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
Tang, H.; Sun, Y.; Wu, D.; Liu, K.; Zhu, J.; and Kang, Z. 2024 · 2024
Closest in time.
Svd-llm: Truncation-aware singular value decomposition for large language model compression
Wang, X.; Zheng, Y.; Wan, Z.; and Zhang, M. 2024 · 2024
Closest in time.
Yang, A.; Yang, B.; Hui, B.; Zheng, B.; Yu, B.; Zhou, C.; Li, C.; Li, C.; Liu, D.; Huang, F.; Dong, G.; Wei, H.; Lin, H.; Tang, J.; Wang, J.; Yang, J.; Tu, J.; Zhang, J.; Ma, J.; Yang, J.; Xu, J.; Zhou, J.; Bai, J.; He, J.; Lin, J.; Dang, K.; Lu, K.; Chen, K.; Yang, K.; Li, M.; Xue, M.; Ni, N.; Zhang, P.; Wang, P.; Peng, R.; Men, R.; Gao, R.; Lin, R.; Wang, S.; Bai, S.; Tan, S.; Zhu, T.; Li, T.; Liu, T.; Ge, W.; Deng, X.; Zhou, X.; Ren, X.; Zhang, X.; Wei, X.; Ren, X.; Liu, X.; Fan, Y.; Yao, Y.; Zhang, Y.; Wan, Y.; Chu, Y.; Liu, Y.; Cui, Z.; Zhang, Z.; Guo, Z.; and Fan, Z. 2024 · 2024
Closest in time.
Exploring post-training quantization in llms from comprehensive study to low rank compensation
Yao, Z.; Wu, X.; Li, C.; Youn, S.; and He, Y. 2024 · 2024
Closest in time.
LQER: Low-Rank Quantization Error Reconstruction for LLMs
Zhang, C.; Cheng, J.; Constantinides, G. A.; and Zhao, Y. 2024 · 2024
Closest in time.