Fetching the paper…
Reading the bibliography…
Despite the impressive performance of LLMs, their widespread adoption faces challenges due to substantial computational and memory requirements during inference.
Pruning versus clipping in neural networks
S. A. Janowsky · 1989
Earlier work this paper cites.
Optimal brain damage
Y. LeCun, J. Denker, and S. Solla · 1989
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
B. Hassibi, D. G Stork, and G. J Wolff · 1993
Earlier work this paper cites.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets
T. N. Sainath, B. Kingsbury, V. Sindhwani, E. Arisoy, and B. Ramabhadran · 2013
Earlier work this paper cites.
Exploiting linear structure within convolutional networks for efficient evaluation
E. Denton, W. Zaremba, J. Bruna, Y. LeCun, and R. Fergus · 2014
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
M. Jaderberg, A. Vedaldi, and A. Zisserman · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Speeding-up convolutional neural networks using fine-tuned cp-decomposition
V. Lebedev, Y. Ganin, M. Rakhuba, I. Oseledets, and V. Lempitsky · 2015
Earlier work this paper cites.
Pruning filters for efficient convnets
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf · 2016
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
P. Molchanov, S. Tyree, T. Karras, T. Aila, and J. Kautz · 2016
Earlier work this paper cites.
Convolutional neural networks with low-rank regularization
C. Tai, T. Xiao, Y. Zhang, X. Wang, and Weinan E · 2016
Earlier work this paper cites.
Learning efficient convolutional networks through network slimming
Z. Liu, J. Li, Z. Shen, G. Huang, S. Yan, and C. Zhang · 2017
Earlier work this paper cites.
Post-training 4-bit quantization of convolution networks for rapid-deployment
R. Banner, Y. Nahshan, E. Hoffer, and D. Soudry · 2019
Earlier work this paper cites.
Low-bit quantization of neural networks for efficient inference
Y. Choukroun, E. Kravchik, F. Yang, and P. Kisilev · 2019
Earlier work this paper cites.
Learned step size quantization
S. K Esser, J. L McKinstry, D. Bablani, R. Appuswamy, and D. S Modha · 2019
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and L. Kristina Toutanova · 2019
Earlier work this paper cites.
Language models are few-shot learners
T Brown, B Mann, N Ryder, M Subbiah, J D Kaplan, P Dhariwal, A Neelakantan, P Shyam, G Sastry, A Askell, et al · 2020
Earlier work this paper cites.
Wrapnet: Neural net inference with ultra-low-resolution arithmetic
R. Ni, H. M. Chu, O. Castañeda, P. Y. Chiang, C. Studer, and T. Goldstein · 2020
Earlier work this paper cites.
Movement pruning: Adaptive sparsity by fine-tuning
V. Sanh, T. Wolf, and A. M. Rush · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J Hu, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al · 2021
Earlier work this paper cites.
Block pruning for faster transformers
F. Lagunas, E. Charlaix, V. Sanh, and A. M. Rush · 2021
Earlier work this paper cites.
Degree-quant: Quantization-aware training for graph neural networks
S. A. Tailor, J. F. Marques, and N. D. Lane · 2021
Cited alongside, same era.
Vision transformer slimming: Multi-dimension searching in continuous optimization space
A Chavan, Z Shen, Z Liu, Z Liu, K T Cheng, and E P Xing · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
T. Dao, D. Fu, S. Ermon, A. Rudra, and Christopher Ré · 2022
Cited alongside, same era.
Llm.int8(): 8-bit matrix multiplication for transformers at scale
T. Dettmers, M. Lewis, Y. Belkada, and L. Zettlemoyer · 2022
Cited alongside, same era.
Compression of generative pre-trained language models via quantization
C. Tao, L. Hou, W. Zhang, L. Shang, X. Jiang, Q. Liu, P. Luo, and N. Wong · 2022
Cited alongside, same era.
Contrastive representation distillation
Yo. Tian, D. Krishnan, and P. Isola · 2022
Fast inference from transformers via speculative decoding
Y. Leviathan, M. Kalman, and Y. Matias · 2023
Later among the works it cites.
Losparse: Structured compression of large language models based on low-rank and sparse approximation
Y. Li, Y. Yu, Q. Zhang, C. Liang, P. He, W. Chen, and T. Zhao · 2023
Later among the works it cites.
Less is more: Task-aware layer-wise distillation for language model compression
C. Liang, S. Zuo, Q. Zhang, P. He, W. Chen, and T. Zhao · 2023
Later among the works it cites.
Awq: Activation-aware weight quantization for llm compression and acceleration
J. Lin, J. Tang, H. Tang, S. Yang, X. Dang, C. Gan, and S. Han · 2023
Later among the works it cites.
Llm-qat: Data-free quantization aware training for large language models
Z. Liu, B. Oguz, C. Zhao, E. Chang, P. Stock, Y. Mehdad, Y. Shi, R. Krishnamoorthi, and V. Chandra · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generalized knowledge distillation for auto-regressive language models
R. Agarwal, N. Vieillard, Y. Zhou, P. Stanczyk, S. Ramos, M. Geist, and O. Bachem · 2023
Cited alongside, same era.
Fluctuation-based adaptive structured pruning for large language models
Y. An, X. Zhao, T. Yu, M. Tang, and J. Wang · 2023
Cited alongside, same era.
Rethinking compression: Reduced order modelling of latent features in large language models
A. Chavan, N. Lele, and D. Gupta · 2023
Cited alongside, same era.
Lorashear: Efficient large language model structured pruning and knowledge recovery
T. Chen, T. Ding, B. Yadav, I. Zharkov, and L. Liang · 2023
Cited alongside, same era.
Disco: Distilling counterfactuals with large language models
Z. Chen, Q. Gao, A. Bosselut, A. Sabharwal, and K. Richardson · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
Cited alongside, same era.
X. Ma, G. Fang, and X. Wang · 2023
Later among the works it cites.
One-Shot Sensitivity-Aware Mixed Sparsity Pruning for Large Language Models
Hang S., Bei L., and Yanmin Q · 2023
Later among the works it cites.
Omniquant: Omnidirectionally calibrated quantization for large language models
W. Shao, M. Chen, Z. Zhang, P. Xu, L. Zhao, Z. Li, K. Zhang, P. Gao, Y. Qiao, and P. Luo · 2023
Later among the works it cites.
The truth is in there: Improving reasoning in language models with layer-selective rank reduction
P. Sharma, J. T. Ash, and D. Misra · 2023
Later among the works it cites.
Powerinfer: Fast large language model serving with a consumer-grade gpu
Y. Song, Z. Mi, H. Xie, and H. Chen · 2023
Later among the works it cites.
A simple and effective pruning approach for large language models
M. Sun, Z. Liu, A. Bair, and J. Z. Kolter · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Later among the works it cites.
Bitnet: Scaling 1-bit transformers for large language models
H. Wang, S. Ma, L. Dong, S. Huang, H. Wang, L. Ma, F. Yang, R. Wang, Y. Wu, and F. Wei · 2023
Later among the works it cites.
Scott: Self-consistent chain-of-thought distillation
P. Wang, Z. Wang, Z. Li, Y. Gao, B. Yin, and X. Ren · 2023
Later among the works it cites.
Zeroquant-fp: A leap forward in llms post-training w4a8 quantization using floating-point formats
X. Wu, Z. Yao, and Y. He · 2023
Later among the works it cites.
Sheared llama: Accelerating language model pre-training via structured pruning
M. Xia, T. Gao, Z. Zeng, and D. Chen · 2023
Later among the works it cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han · 2023
Later among the works it cites.
Tensorgpt: Efficient compression of the embedding layer in llms based on the tensor-train decomposition
M. Xu, Y. Lei Xu, and D. P. Mandic · 2023
Later among the works it cites.
Loraprune: Pruning meets low-rank parameter-efficient fine-tuning
M. Zhang, H. Chen, C. Shen, Z. Yang, L. Ou, X. Yu, and B. Zhuang · 2023
Later among the works it cites.
A survey on model compression for large language models
X. Zhu, J. Li, Y. Liu, C. Ma, and W. Wang · 2023
Later among the works it cites.