Fetching the paper…
Reading the bibliography…
The remarkable success of Large Language Models (LLMs) relies heavily on their substantial scale, which poses significant challenges during model deployment in terms of latency and memory consumption.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
On information and sufficiency
Kullback, S.; and Leibler, R. A. 1951 · 1951
Earlier work this paper cites.
Optimal brain damage
LeCun, Y.; Denker, J.; and Solla, S. 1990 · 1990
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
Hassibi, B.; Stork, D. G.; and Wolff, G. J. 1993 · 1993
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X.; and Bengio, Y. 2010 · 2010
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y.; Léonard, N.; and Courville, A. 2013 · 2013
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Han, S.; Pool, J.; Tran, J.; and Dally, W. J. 2015 · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Han, S.; Mao, H.; and Dally, W. J. 2016 · 2016
Earlier work this paper cites.
Learning efficient convolutional networks through network slimming
Liu, Z.; Li, J.; Shen, Z.; Huang, G.; Yan, S.; and Zhang, C. 2017 · 2017
Earlier work this paper cites.
Think you have solved question answering? Try arc, the ai2 reasoning challenge
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T.; Clark, P.; Khot, T.; and Sabharwal, A. 2018 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2018 · 2018
Earlier work this paper cites.
To prune, or not to prune: Exploring the efficacy of pruning for model compression
Zhu, M.; and Gupta, S. 2018 · 2018
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural Yes/No questions
Clark, C.; Lee, K.; Chang, M.-W.; Kwiatkowski, T.; Collins, M.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J.; and Carbin, M. 2019 · 2019
Earlier work this paper cites.
SNIP: Single-shot network pruning based on connection sensitivity
Lee, N.; Ajanthan, T.; and Torr, P. H. 2019 · 2019
Earlier work this paper cites.
Importance estimation for neural network pruning
Molchanov, P.; Mallya, A.; Tyree, S.; Frosio, I.; and Kautz, J. 2019 · 2019
Cited alongside, same era.
Structured pruning of large language models
Wang, Z.; Wohlwend, J.; and Lei, T. 2019 · 2019
Cited alongside, same era.
HellaSwag: Can a machine really finish your sentence?
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 2019
Cited alongside, same era.
Rigging the lottery: Making all tickets winners
Evci, U.; Gale, T.; Menick, J.; Castro, P. S.; and Elsen, E. 2020 · 2020
Cited alongside, same era.
Tinybert: Distilling bert for natural language understanding
Jiao, X.; Yin, Y.; Shang, L.; Jiang, X.; Chen, X.; Li, L.; Wang, F.; and Liu, Q. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Emergent abilities of large language models
Wei, J.; Tay, Y.; Bommasani, R.; Raffel, C.; Zoph, B.; Borgeaud, S.; Yogatama, D.; Bosma, M.; Zhou, D.; Metzler, D.; Chi, E. H.; Hashimoto, T.; Vinyals, O.; Liang, P.; Dean, J.; and Fedus, W. 2022 · 2022
Later among the works it cites.
OPT: Open Pre-trained Transformer Language Models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; Mihaylov, T.; Ott, M.; Shleifer, S.; Shuster, K.; Simig, D.; Koura, P. S.; Sridhar, A.; Wang, T.; and Zettlemoyer, L. 2022 · 2022
Later among the works it cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Later among the works it cites.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Frantar, E.; and Alistarh, D. 2023 · 2023
Later among the works it cites.
Sparse finetuning for inference acceleration of large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Comparing rewinding and fine-tuning in neural network pruning
Renda, A.; Frankle, J.; and Carbin, M. 2020 · 2020
Cited alongside, same era.
Woodfisher: Efficient second-order approximation for neural network compression
Singh, S. P.; and Alistarh, D. 2020 · 2020
Cited alongside, same era.
Mobilebert: A compact task-agnostic bert for resource-limited devices
Sun, Z.; Yu, H.; Song, X.; Liu, R.; Yang, Y.; and Zhou, D. 2020 · 2020
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Cited alongside, same era.
A framework for few-shot language model evaluation
Gao, L.; Tow, J.; Biderman, S.; Black, S.; DiPofi, A.; Foster, C.; Golding, L.; Hsu, J.; McDonell, K.; Muennighoff, N.; et al. 2021 · 2021
Cited alongside, same era.
Accelerated sparse neural training: A provable and efficient method to find n: m transposable masks
Hubara, I.; Chmiel, B.; Island, M.; Banner, R.; Naor, J.; and Soudry, D. 2021 · 2021
Cited alongside, same era.
Kurtic, E.; Kuznedelev, D.; Frantar, E.; Goin, M.; and Alistarh, D. 2023 · 2023
Later among the works it cites.
Gradient-free structured pruning with unlabeled data
Nova, A.; Dai, H.; and Schuurmans, D. 2023 · 2023
Later among the works it cites.
Unmasking the lottery ticket hypothesis: What’s encoded in a winning ticket’s mask?
Paul, M.; Chen, F.; Larsen, B. W.; Frankle, J.; Ganguli, S.; and Dziugaite, G. K. 2023 · 2023
Later among the works it cites.
Upop: Unified and progressive pruning for compressing vision-language transformers
Shi, D.; Tao, C.; Jin, Y.; Yang, Z.; Yuan, C.; and Wang, J. 2023 · 2023
Later among the works it cites.
A simple and effective pruning approach for large language models
Sun, M.; Liu, Z.; Bair, A.; and Kolter, J. Z. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Later among the works it cites.
A three-regime model of network pruning
Zhou, Y.; Yang, Y.; Chang, A.; and Mahoney, M. W. 2023 · 2023
Later among the works it cites.
MiniLLM: Knowledge distillation of large language models
Gu, Y.; Dong, L.; Wei, F.; and Huang, M. 2024 · 2024
Closest in time.
Compressing large language models by joint sparsification and quantization
Guo, J.; Wu, J.; Wang, Z.; Liu, J.; Yang, G.; Ding, Y.; Gong, R.; Qin, H.; and Liu, X. 2024 · 2024
Closest in time.
Accelerating transformer pre-training with 2:4 sparsity
Hu, Y.; Zhao, K.; Huang, W.; Chen, J.; and Zhu, J. 2024 · 2024
Closest in time.
Awq: Activation-aware weight quantization for LLM compression and acceleration
Lin, J.; Tang, J.; Tang, H.; Yang, S.; Dang, X.; and Han, S. 2024 · 2024
Closest in time.
Compressing LLMs: The truth is rarely pure and never simple
Nowak, A. I.; Grooten, B.; Mocanu, D. C.; and Tabor, J. 2024 · 2024
Closest in time.
Sheared llama: Accelerating language model pre-training via structured pruning
Xia, M.; Gao, T.; Zeng, Z.; and Chen, D. 2024 · 2024
Closest in time.