Fetching the paper…
Reading the bibliography…
Determining the optimal size of a neural network is critical, as it directly impacts runtime performance and memory usage.
Fluctuation-Based Adaptive Structured Pruning for Large Language Models
An, Y.; Zhao, X.; Yu, T.; Tang, M.; and Wang, J. 2024 · 2014
Earlier work this paper cites.
A Kernel Independence Test for Random Processes
Chwialkowski, K.; and Gretton, A. 2014 · 2014
Earlier work this paper cites.
Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding
Han, S.; Mao, H.; and Dally, W. J. 2016 · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation
Liu, C.; Lowe, R.; Serban, I.; Noseworthy, M.; Charlin, L.; and Pineau, J. 2016 · 2016
Earlier work this paper cites.
Pruning Convolutional Neural Networks for Resource Efficient Transfer Learning
Molchanov, P.; Tyree, S.; Karras, T.; Aila, T.; and Kautz, J. 2016 · 2016
Earlier work this paper cites.
Channel Pruning for Accelerating Very Deep Neural Networks
He, Y.; Zhang, X.; and Sun, J. 2017 · 2017
Earlier work this paper cites.
Pruning Filters for Efficient ConvNets
Li, H.; Kadav, A.; Durdanovic, I.; Samet, H.; and Graf, H. P. 2017 · 2017
Earlier work this paper cites.
Learning Efficient Convolutional Networks through Network Slimming
Liu, Z.; Li, J.; Shen, Z.; Huang, G.; Yan, S.; and Zhang, C. 2017 · 2017
Earlier work this paper cites.
ThiNet: A Filter Level Pruning Method for Deep Neural Network Compression
Luo, J.; Wu, J.; and Lin, W. 2017 · 2017
Earlier work this paper cites.
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
State-of-the-art in artificial neural network applications: A survey
Abiodun, O. I.; Jantan, A.; Omolara, A. E.; Dada, K. V.; Mohamed, N. A.; and Arshad, H. 2018 · 2018
Earlier work this paper cites.
Data-Driven Sparse Structure Selection for Deep Neural Networks
Huang, Z.; and Wang, N. 2018 · 2018
Earlier work this paper cites.
Similarity of Neural Network Representations Revisited
Kornblith, S.; Norouzi, M.; Lee, H.; and Hinton, G. E. 2019 · 2019
Earlier work this paper cites.
Revealing the Dark Secrets of BERT
Kovaleva, O.; Romanov, A.; Rogers, A.; and Rumshisky, A. 2019 · 2019
Earlier work this paper cites.
Towards Optimal Structured CNN Pruning via Generative Adversarial Learning
Lin, S.; Ji, R.; Yan, C.; Zhang, B.; Cao, L.; Ye, Q.; Huang, F.; and Doermann, D. S. 2019 · 2019
Earlier work this paper cites.
Are Sixteen Heads Really Better than One?
Michel, P.; Levy, O.; and Neubig, G. 2019 · 2019
Cited alongside, same era.
Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
Voita, E.; Talbot, D.; Moiseev, F.; Sennrich, R.; and Titov, I. 2019 · 2019
Cited alongside, same era.
MINT: Deep Network Compression via Mutual Information-based Neuron Trimming
Ganesh, M. R.; Corso, J. J.; and Sekeh, S. Y. 2020 · 2020
Cited alongside, same era.
A survey of the usages of deep learning for natural language processing
Otter, D. W.; Medina, J. R.; and Kalita, J. K. 2020 · 2020
Cited alongside, same era.
How fine can fine-tuning be? Learning efficient language models
Radiya-Dixit, E.; and Wang, X. 2020 · 2020
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
SparseGPT: Massive Language Models Can be Accurately Pruned in One-Shot
Frantar, E.; and Alistarh, D. 2023 · 2023
Later among the works it cites.
Compresso: Structured Pruning with Collaborative Prompting Learns Compact Large Language Models
Guo, S.; Xu, J.; Zhang, L. L.; and Yang, M. 2023 · 2023
Later among the works it cites.
LLM-Pruner: On the Structural Pruning of Large Language Models
Ma, X.; Fang, G.; and Wang, X. 2023 · 2023
Later among the works it cites.
DeepArc: Modularizing Neural Networks for the Model Maintenance
Ren, X.; Lin, Y.; Xue, Y.; Liu, R.; Sun, J.; Feng, Z.; and Dong, J. S. 2023 · 2023
Later among the works it cites.
Transformers in medical imaging: A survey
Shamshad, F.; Khan, S. H.; Zamir, S. W.; Khan, M. H.; Hayat, M.; Khan, F. S.; and Fu, H. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Layer-wise Model Pruning based on Mutual Information
Fan, C.; Li, J.; Zhang, T.; Ao, X.; Wu, F.; Meng, Y.; and Sun, X. 2021 · 2021
Cited alongside, same era.
Block Pruning For Faster Transformers
Lagunas, F.; Charlaix, E.; Sanh, V.; and Rush, A. M. 2021 · 2021
Cited alongside, same era.
Transformers in Vision: A Survey
Khan, S. H.; Naseer, M.; Hayat, M.; Zamir, S. W.; Khan, F. S.; and Shah, M. 2022 · 2022
Cited alongside, same era.
A survey of transformers
Lin, T.; Wang, Y.; Liu, X.; and Qiu, X. 2022 · 2022
Cited alongside, same era.
Scale Efficiently: Insights from Pretraining and Finetuning Transformers
Tay, Y.; Dehghani, M.; Rao, J.; Fedus, W.; Abnar, S.; Chung, H. W.; Narang, S.; Yogatama, D.; Vaswani, A.; and Metzler, D. 2022 · 2022
Cited alongside, same era.
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
Zaken, E. B.; Goldberg, Y.; and Ravfogel, S. 2022 · 2022
Cited alongside, same era.
Wang, H.; Ge, C.; Chen, H.; and Sun, X. 2023a · 2023
Later among the works it cites.
Maximizing Spatio-Temporal Entropy of Deep 3D CNNs for Efficient Video Recognition
Wang, J.; Sun, Z.; Qian, Y.; Gong, D.; Sun, X.; Lin, M. C.; Pagnucco, M.; and Song, Y. 2023b · 2023
Later among the works it cites.
SliceGPT: Compress Large Language Models by Deleting Rows and Columns
Ashkboos, S.; Croci, M. L.; Nascimento, M. G. D.; Hoefler, T.; and Hensman, J. 2024 · 2024
Closest in time.
Not all Layers of LLMs are Necessary during Inference
Fan, S.; Jiang, X.; Li, X.; Meng, X.; Han, P.; Shang, S.; Sun, A.; Wang, Y.; and Wang, Z. 2024 · 2024
Closest in time.
Structured Pruning for Deep Convolutional Neural Networks: A Survey
He, Y.; and Xiao, L. 2024 · 2024
Closest in time.
Shortened LLaMA: A Simple Depth Pruning for Large Language Models
Kim, B.; Kim, G.; Kim, T.; Castells, T.; Choi, S.; Shin, J.; and Song, H. 2024 · 2024
Closest in time.
Pruning via Merging: Compressing LLMs via Manifold Alignment Based Layer Merging
Liu, D.; Qin, Z.; Wang, H.; Yang, Z.; Wang, Z.; Rong, F.; Liu, Q.; Hao, Y.; Chen, X.; Fan, C.; Lv, Z.; Tu, Z.; Chu, D.; and Sui, D. 2024 · 2024
Closest in time.
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Men, X.; Xu, M.; Zhang, Q.; Wang, B.; Lin, H.; Lu, Y.; Han, X.; and Chen, W. 2024 · 2024
Closest in time.
A Simple and Effective Pruning Approach for Large Language Models
Sun, M.; Liu, Z.; Bair, A.; and Kolter, J. Z. 2024 · 2024
Closest in time.
LaCo: Large Language Model Pruning via Layer Collapse
Yang, Y.; Cao, Z.; and Zhao, H. 2024 · 2024
Closest in time.