Fetching the paper…
Reading the bibliography…
Layer pruning has emerged as a promising technique for compressing large language models (LLMs) while achieving acceleration proportional to the pruning ratio.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Clark, C.; Lee, K.; Chang, M.-W.; Kwiatkowski, T.; Collins, M.; and Toutanova, K. 2019 · 1905
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 1905
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Clark, C.; Lee, K.; Chang, M.-W.; Kwiatkowski, T.; Collins, M.; and Toutanova, K. 2019 · 1905
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 1905
Earlier work this paper cites.
Building a Large Annotated Corpus of English: The Penn Treebank
Marcus, M. P.; Santorini, B.; and Marcinkiewicz, M. A. 1993 · 1993
Earlier work this paper cites.
Building a Large Annotated Corpus of English: The Penn Treebank
Marcus, M. P.; Santorini, B.; and Marcinkiewicz, M. A. 1993 · 1993
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020 · 2009
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020 · 2009
Earlier work this paper cites.
The winograd schema challenge
Levesque, H.; Davis, E.; and Morgenstern, L. 2012 · 2012
Earlier work this paper cites.
The winograd schema challenge
Levesque, H.; Davis, E.; and Morgenstern, L. 2012 · 2012
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2016 · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2016 · 2016
Earlier work this paper cites.
Race: Large-scale reading comprehension dataset from examinations
Lai, G.; Xie, Q.; Liu, H.; Yang, Y.; and Hovy, E. 2017 · 2017
Earlier work this paper cites.
Race: Large-scale reading comprehension dataset from examinations
Lai, G.; Xie, Q.; Liu, H.; Yang, Y.; and Hovy, E. 2017 · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018 · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018 · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y.; Zellers, R.; Gao, J.; Choi, Y.; et al. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K.; Le Bras, R.; Bhagavatula, C.; and Choi, Y. 2020 · 2020
Earlier work this paper cites.
Superglue: Learning feature matching with graph neural networks
Sarlin, P.-E.; DeTone, D.; Malisiewicz, T.; and Rabinovich, A. 2020 · 2020
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y.; Zellers, R.; Gao, J.; Choi, Y.; et al. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K.; Le Bras, R.; Bhagavatula, C.; and Choi, Y. 2020 · 2020
Earlier work this paper cites.
Superglue: Learning feature matching with graph neural networks
Sarlin, P.-E.; DeTone, D.; Malisiewicz, T.; and Rabinovich, A. 2020 · 2020
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. 2023 · 2023
Cited alongside, same era.
Llm-pruner: On the structural pruning of large language models
Ma, X.; Fang, G.; and Wang, X. 2023 · 2023
Cited alongside, same era.
A simple and effective pruning approach for large language models
Sun, M.; Liu, Z.; Bair, A.; and Kolter, J. Z. 2023 · 2023
Cited alongside, same era.
LLaMA-NAS: Efficient Neural Architecture Search for Large Language Models
Sarah, A.; Sridhar, S. N.; Szankin, M.; and Sundaresan, S. 2024 · 2024
Later among the works it cites.
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
Song, J.; Oh, K.; Kim, T.; Kim, H.; Kim, Y.; and Kim, J.-J. 2024 · 2024
Later among the works it cites.
Llm pruning and distillation in practice: The minitron approach
Sreenivas, S. T.; Muralidharan, S.; Joshi, R.; Chochowski, M.; Mahabaleshwarkar, A. S.; Shen, G.; Zeng, J.; Chen, Z.; Suhara, Y.; Diao, S.; et al. 2024 · 2024
Later among the works it cites.
Flatquant: Flatness matters for llm quantization
Sun, Y.; Liu, R.; Bai, H.; Bao, H.; Zhao, K.; Li, Y.; Hu, J.; Yu, X.; Hou, L.; Yuan, C.; et al. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Cited alongside, same era.
van der Ouderaa, T. F.; Nagel, M.; Van Baalen, M.; Asano, Y. M.; and Blankevoort, T. 2023 · 2023
Cited alongside, same era.
Sheared llama: Accelerating language model pre-training via structured pruning
Xia, M.; Gao, T.; Zeng, Z.; and Chen, D. 2023 · 2023
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. 2023 · 2023
Cited alongside, same era.
Llm-pruner: On the structural pruning of large language models
Ma, X.; Fang, G.; and Wang, X. 2023 · 2023
Cited alongside, same era.
A simple and effective pruning approach for large language models
Sun, M.; Liu, Z.; Bair, A.; and Kolter, J. Z. 2023 · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Cited alongside, same era.
Fluctuation-based adaptive structured pruning for large language models
An, Y.; Zhao, X.; Yu, T.; Tang, M.; and Wang, J. 2024 · 2024
Later among the works it cites.
Compressing large language models by streamlining the unimportant layer
Chen, X.; Hu, Y.; and Zhang, J. 2024 · 2024
Later among the works it cites.
Streamlining redundant layers to compress large language models
Chen, X.; Hu, Y.; Zhang, J.; Wang, Y.; Li, C.; and Chen, H. 2024 · 2024
Later among the works it cites.
Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024 · 2024
Later among the works it cites.
The Unreasonable Ineffectiveness of the Deeper Layers
Gromov, A.; Tirumala, K.; Shapourian, H.; Glorioso, P.; and Roberts, D. 2024 · 2024
Later among the works it cites.
SP3: Enhancing Structured Pruning via PCA Projection
Hu, Y.; Zhang, J.; Zhao, Z.; Zhao, C.; Chen, X.; Li, C.; and Chen, H. 2024 · 2024
Later among the works it cites.
Shortened llama: A simple depth pruning for large language models
Kim, B.-K.; Kim, G.; Kim, T.-H.; Castells, T.; Choi, S.; Shin, J.; and Song, H.-K. 2024 · 2024
Later among the works it cites.
Liu, A.; Feng, B.; Xue, B.; Wang, B.; Wu, B.; Lu, C.; Zhao, C.; Deng, C.; Zhang, C.; Ruan, C.; et al. 2024 · 2024
Later among the works it cites.
Shortgpt: Layers in large language models are more redundant than you expect
Men, X.; Xu, M.; Zhang, Q.; Wang, B.; Lin, H.; Lu, Y.; Han, X.; and Chen, W. 2024 · 2024
Later among the works it cites.
Compact language models via pruning and knowledge distillation
Muralidharan, S.; Turuvekere Sreenivas, S.; Joshi, R.; Chochowski, M.; Patwary, M.; Shoeybi, M.; Catanzaro, B.; Kautz, J.; and Molchanov, P. 2024 · 2024
Later among the works it cites.
LLaMA-NAS: Efficient Neural Architecture Search for Large Language Models
Sarah, A.; Sridhar, S. N.; Szankin, M.; and Sundaresan, S. 2024 · 2024
Later among the works it cites.
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
Song, J.; Oh, K.; Kim, T.; Kim, H.; Kim, Y.; and Kim, J.-J. 2024 · 2024
Later among the works it cites.
Llm pruning and distillation in practice: The minitron approach
Sreenivas, S. T.; Muralidharan, S.; Joshi, R.; Chochowski, M.; Mahabaleshwarkar, A. S.; Shen, G.; Zeng, J.; Chen, Z.; Suhara, Y.; Diao, S.; et al. 2024 · 2024
Later among the works it cites.
Flatquant: Flatness matters for llm quantization
Sun, Y.; Liu, R.; Bai, H.; Bao, H.; Zhao, K.; Li, Y.; Hu, J.; Yu, X.; Hou, L.; Yuan, C.; et al. 2024 · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P.; Bi, X.; et al. 2025 · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Team, K.; Du, A.; Gao, B.; Xing, B.; Jiang, C.; Chen, C.; Li, C.; Xiao, C.; Du, C.; Liao, C.; et al. 2025 · 2025
Closest in time.
Team, Q. 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P.; Bi, X.; et al. 2025 · 2025
Closest in time.
Kimi k1. 5: Scaling reinforcement learning with llms
Team, K.; Du, A.; Gao, B.; Xing, B.; Jiang, C.; Chen, C.; Li, C.; Xiao, C.; Du, C.; Liao, C.; et al. 2025 · 2025
Closest in time.
Team, Q. 2025 · 2025
Closest in time.