Fetching the paper…
Reading the bibliography…
Conventional wisdom in pruning Transformer-based language models is that pruning reduces the model expressiveness and thus is more likely to underfit rather than overfit.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Reweighted proximal pruning for large-scale language representation
Fu-Ming Guo, Sijia Liu, Finlay S Mungall, Xue Lin, and Yanzhi Wang. 2019 · 1909
Earlier work this paper cites.
Exponentially small bounds on the expected optimum of the partition and subset sum problems
George S Lueker. 1998 · 1998
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Earlier work this paper cites.
To prune, or not to prune: Exploring the efficacy of pruning for model compression
Michael H Zhu and Suyog Gupta. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2019 · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 2019
Earlier work this paper cites.
An overview of overfitting and its solutions
Xue Ying. 2019 · 2019
Cited alongside, same era.
Few shot network compression via cross distillation
Haoli Bai, Jiaxiang Wu, Irwin King, and Michael Lyu. 2020 · 2020
Cited alongside, same era.
What is the state of neural network pruning?
Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag. 2020 · 2020
Cited alongside, same era.
Edge camera system using deep learning method with model compression on embedded applications
Yun Won Choi and Jang Woon Baek. 2020 · 2020
Cited alongside, same era.
Sparsity through evolutionary pruning prevents neuronal networks from overfitting
Richard C Gerum, André Erpenbeck, Patrick Krauss, and Achim Schilling. 2020 · 2020
Cited alongside, same era.
Compressing bert: Studying the effects of weight pruning on transfer learning
Mitchell Gordon, Kevin Duh, and Nicholas Andrews. 2020 · 2020
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020 · 2020
Later among the works it cites.
BERT-of-theseus: Compressing BERT by progressive module replacing
Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou. 2020 · 2020
Later among the works it cites.
Et: re-thinking self-attention for transformer models on gpus
Shiyang Chen, Shaoyi Huang, Santosh Pandey, Bingbing Li, Guang R Gao, Long Zheng, Caiwen Ding, and Hang Liu. 2021 · 2021
Closest in time.
Hmc-tran: A tensor-core inspired hierarchical model compression for transformer-based dnns on gpu
Shaoyi Huang, Shiyang Chen, Hongwu Peng, Daniel Manu, Zhenglun Kong, Geng Yuan, Lei Yang, Shusen Wang, Hang Liu, and Caiwen Ding. 2021 · 2021
Closest in time.
Npas: A compiler-aware framework of unified network pruning and architecture search for beyond real-time mobile acceleration
Zhengang Li, Geng Yuan, Wei Niu, Pu Zhao, Yanyu Li, Yuxuan Cai, Xuan Shen, Zheng Zhan, Zhenglun Kong, Qing Jin, et al. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Cited alongside, same era.
Improved knowledge distillation via teacher assistant
Seyed Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh. 2020 · 2020
Cited alongside, same era.
Optimal lottery tickets via subset sum: Logarithmic over-parameterization is sufficient
Ankit Pensia, Shashank Rajput, Alliot Nagle, Harit Vishwakarma, and Dimitris Papailiopoulos. 2020 · 2020
Cited alongside, same era.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2020
Cited alongside, same era.
Meta-learning with network pruning
Hongduan Tian, Bo Liu, Xiao-Tong Yuan, and Qingshan Liu. 2020 · 2020
Cited alongside, same era.
Towards extremely compact rnns for video recognition with fully decomposed hierarchical tucker structure
Miao Yin, Siyu Liao, Xiao-Yang Liu, Xiaodong Wang, and Bo Yuan. 2021a
Cited in the paper.
Closest in time.
Accelerating transformer-based deep learning models on fpgas using column balanced block pruning
Hongwu Peng, Shaoyi Huang, Tong Geng, Ang Li, Weiwen Jiang, Hang Liu, Shusen Wang, and Caiwen Ding. 2021 · 2021
Closest in time.
Accelerating framework of transformer by hardware design and model compression co-optimization
Panjie Qi, Edwin Hsing-Mean Sha, Qingfeng Zhuge, Hongwu Peng, Shaoyi Huang, Zhenglun Kong, Yuhong Song, and Bingbing Li. 2021 · 2021
Closest in time.
Progressive network grafting for few-shot knowledge distillation
Chengchao Shen, Xinchao Wang, Youtan Yin, Jie Song, Sihui Luo, and Mingli Song. 2021 · 2021
Closest in time.
Against membership inference attack: Pruning is all you need
Yijue Wang, Chenghong Wang, Zigeng Wang, Shanglin Zhou, Hang Liu, Jinbo Bi, Caiwen Ding, and Sanguthevar Rajasekaran. 2021 · 2021
Closest in time.
Rethinking network pruning – under the pre-train and fine-tune paradigm
Dongkuan Xu, Ian En-Hsu Yen, Jinxi Zhao, and Zhibin Xiao. 2021 · 2021
Closest in time.