Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have achieved remarkable success in natural language processing tasks, but their massive size and computational demands hinder their deployment in resource-constrained environments.
LeCun, Y., Denker, J.S., Solla, S.A.: Optimal brain damage. In: Proceedings of NeurIPS. pp. 598–605 (1989)
1989
Earlier work this paper cites.
Merity, S., et al.: Pointer sentinel mixture models. In: Proceedings of ICLR (2017)
2017
Earlier work this paper cites.
Jiang, C., et al.: Efficient DNN neuron pruning by minimizing layer-wise nonlinear reconstruction error. In: Proceedings of IJCAI. pp. 2298–2304 (2018)
2018
Earlier work this paper cites.
Devlin, J., et al.: BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of NAACL-HLT. pp. 4171–4186 (2019)
2019
Earlier work this paper cites.
Voita, E., et al.: Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. In: Proceedings of ACL. pp. 5797–5808 (2019)
2019
Earlier work this paper cites.
Raffel, C., et al.: Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21
2020
Earlier work this paper cites.
Wang, Z.: Sparsert: Accelerating unstructured sparsity on gpus for deep learning inference. In: Proceedings of PACT. pp. 31–42 (2020)
2020
Earlier work this paper cites.
Xin, J., et al.: Deebert: Dynamic early exiting for accelerating BERT inference. In: Proceedings of ACL. pp. 2246–2251 (2020)
2020
Earlier work this paper cites.
Zhou, A., et al.: Learning N: M fine-grained structured sparse neural networks from scratch. In: Proceedings of ICLR (2021)
2021
Earlier work this paper cites.
Scao, T.L., et al.: BLOOM: A 176b-parameter open-access multilingual language model. CoRR (2022)
2022
Cited alongside, same era.
Thoppilan, R., et al.: Lamda: Language models for dialog applications. CoRR (2022)
2022
Cited alongside, same era.
Zhang, S., et al.: OPT: open pre-trained transformer language models. CoRR (2022)
2022
Cited alongside, same era.
Zhang, Y., et al.: Learning best combination for efficient N: M sparsity. In: Proceedings of NeurIPS (2022)
2022
Cited alongside, same era.
Chowdhery, A., et al.: Palm: Scaling language modeling with pathways. Journal of Machine Learning Research 24
2023
Cited alongside, same era.
Frantar, E., Alistarh, D.: Sparsegpt: Massive language models can be accurately pruned in one-shot. In: Proceedings of ICML (2023)
Zeng, A., et al.: GLM-130B: an open bilingual pre-trained model. In: Proceedings of ICLR (2023)
2023
Later among the works it cites.
Ashkboos, S., et al.: Slicegpt: Compress large language models by deleting rows and columns. In: Proceedings of ICLR (2024)
2024
Later among the works it cites.
Gao, L., et al.: A framework for few-shot language model evaluation (07 2024)
2024
Later among the works it cites.
Huang, L., et al.: RAEE: A training-free retrieval-augmented early exiting framework for efficient inference. CoRR (2024)
2024
Later among the works it cites.
Song, J., et al.: SLEB: streamlining llms through redundancy verification and elimination of transformer blocks. In: Proceedings of ICML (2024)
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Ma, X., Fang, G., Wang, X.: Llm-pruner: On the structural pruning of large language models. In: Proceedings of NeurIPS (2023)
2023
Cited alongside, same era.
Touvron, H., et al.: Llama 2: Open foundation and fine-tuned chat models. CoRR (2023)
2023
Cited alongside, same era.
Touvron, H., et al.: Llama: Open and efficient foundation language models. CoRR (2023)
2023
Cited alongside, same era.
Sun, M., et al.: A simple and effective pruning approach for large language models. In: Proceedings of ICLR (2024)
2024
Later among the works it cites.
Zhang, P., et al.: Tinyllama: An open-source small language model. CoRR (2024)
2024
Later among the works it cites.
Zhang, Y., et al.: Dynamic sparse no training: Training-free fine-tuning for sparse llms. In: Proceedings of ICLR (2024)
2024
Later among the works it cites.