Fetching the paper…
Reading the bibliography…
Model quantification uses low bit-width values to represent the weight matrices of existing models to be quantized, which is a promising approach to reduce both storage and computational overheads of deploying highly anticipated LLMs.
Sulle funzioni bilineari, giomale di mathematiche ad uso studenti delle uninersita. 11, 98–106.(an english translation by d boley is available as university of minnesota, department of computer science)
E. Beltrami · 1990
Earlier work this paper cites.
Positive matrix factorization: A non-negative factor model with optimal utilization of error estimates of data values
P. Paatero and U. Tapper · 1994
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Think you have solved question answering? Try ARC, the AI2 Reasoning Challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
C. Clark, K. Lee, M.-W. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Patient knowledge distillation for BERT model compression
S. Sun, Y. Cheng, Z. Gan, and J. Liu · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
PIQA: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, J. Gao, Y. Choi, et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Compressing pre-trained language models by matrix decomposition
M. B. Noach and Y. Goldberg · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2021
Cited alongside, same era.
LLM.int8(): 8-bit matrix multiplication for transformers at scale
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer · 2022
Cited alongside, same era.
GPTQ: Accurate post-training quantization for generative pre-trained transformers
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh · 2022
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
S. Biderman, H. Schoelkopf, Q. G. Anthony, H. Bradley, K. O’Brien, E. Hallahan, M. A. Khan, S. Purohit, U. S. Prashanth, E. Raff, et al · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
BitNet: Scaling 1-bit transformers for large language models
H. Wang, S. Ma, L. Dong, S. Huang, H. Wang, L. Ma, F. Yang, R. Wang, Y. Wu, and F. Wei · 2023
Later among the works it cites.
SmoothQuant: Accurate and efficient post-training quantization for large language models
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han · 2023
Later among the works it cites.
M. Xu, Y. L. Xu, and D. P. Mandic · 2023
Later among the works it cites.
A survey on model compression for large language models
X. Zhu, J. Li, Y. Liu, C. Ma, and W. Wang · 2023
Later among the works it cites.
On-policy distillation of language models: Learning from self-generated mistakes
R. Agarwal, N. Vieillard, Y. Zhou, P. Stanczyk, S. R. Garea, M. Geist, and O. Bachem · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The case for 4-bit precision: k-bit inference scaling laws
T. Dettmers and L. Zettlemoyer · 2023
Cited alongside, same era.
QLoRA: Efficient finetuning of quantized LLMs
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer · 2023
Cited alongside, same era.
SparseGPT: Massive language models can be accurately pruned in one-shot
E. Frantar and D. Alistarh · 2023
Cited alongside, same era.
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
C.-Y. Hsieh, C.-L. Li, C.-k. Yeh, H. Nakhost, Y. Fujii, A. Ratner, R. Krishna, C.-Y. Lee, and T. Pfister · 2023
Cited alongside, same era.
Memory-efficient fine-tuning of compressed large language models via sub-4-bit integer quantization
J. Kim, J. H. Lee, S. Kim, J. Park, K. M. Yoo, S. J. Kwon, and D. Lee · 2023
Cited alongside, same era.
LLM-Pruner: On the structural pruning of large language models
X. Ma, G. Fang, and X. Wang · 2023
Cited alongside, same era.
B. Peng, C. Li, P. He, M. Galley, and J. Gao · 2023
Cited alongside, same era.
Closest in time.
SpQR: A sparse-quantized representation for near-lossless llm weight compression
T. Dettmers, R. Svirschevski, V. Egiazarian, D. Kuznedelev, E. Frantar, S. Ashkboos, A. Borzunov, T. Hoefler, and D. Alistarh · 2024
Closest in time.
SqueezeLLM: Dense-and-sparse quantization
S. Kim, C. Hooper, A. Gholami, Z. Dong, X. Li, S. Shen, M. W. Mahoney, and K. Keutzer · 2024
Closest in time.
AWQ: Activation-aware weight quantization for on-device llm compression and acceleration
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han · 2024
Closest in time.
LLM-QAT: Data-free quantization aware training for large language models
Z. Liu, B. Oguz, C. Zhao, E. Chang, P. Stock, Y. Mehdad, Y. Shi, R. Krishnamoorthi, and V. Chandra · 2024
Closest in time.
OmniQuant: Omnidirectionally calibrated quantization for large language models
W. Shao, M. Chen, Z. Zhang, P. Xu, L. Zhao, Z. Li, K. Zhang, P. Gao, Y. Qiao, and P. Luo · 2024
Closest in time.
A simple and effective pruning approach for large language models
M. Sun, Z. Liu, A. Bair, and J. Z. Kolter · 2024
Closest in time.
QuIP#: Even better llm quantization with hadamard incoherence and lattice codebooks
A. Tseng, J. Chee, Q. Sun, V. Kuleshov, and C. De Sa · 2024
Closest in time.
Efficient large language models: A survey
Z. Wan, X. Wang, C. Liu, S. Alam, Y. Zheng, J. Liu, Z. Qu, S. Yan, Y. Zhu, Q. Zhang, M. Chowdhury, and M. Zhang · 2024
Closest in time.
TinyLlama: An open-source small language model
P. Zhang, G. Zeng, T. Wang, and W. Lu · 2024
Closest in time.