Fetching the paper…
Reading the bibliography…
Auto-regressive large language models such as GPT-3 require enormous computational resources to use.
The state of sparsity in deep neural networks, 2019
Trevor Gale, Erich Elsen, and Sara Hooker · 1902
Earlier work this paper cites.
Language models as knowledge bases?, 2019
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel · 1909
Earlier work this paper cites.
Dialogpt: Large-scale generative pre-training for conversational response generation, 2019
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan · 1911
Earlier work this paper cites.
Skeletonization: A technique for trimming the fat from a network via relevance assessment
Michael C Mozer and Paul Smolensky · 1988
Earlier work this paper cites.
Pruning versus clipping in neural networks
Steven A. Janowsky · 1989
Earlier work this paper cites.
A simple procedure for pruning back-propagation trained neural networks
E.D. Karnin · 1990
Earlier work this paper cites.
Pruning algorithms-a survey
Russell Reed · 1993
Earlier work this paper cites.
Learned threshold pruning, 2020
Kambiz Azarian, Yash Bhalgat, Jinwon Lee, and Tijmen Blankevoort · 2003
Earlier work this paper cites.
What is the state of neural network pruning?, 2020
Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag · 2003
Earlier work this paper cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2005
Earlier work this paper cites.
Movement pruning: Adaptive sparsity by fine-tuning, 2020
Victor Sanh, Thomas Wolf, and Alexander M. Rush · 2005
Earlier work this paper cites.
Bingbing Li, Zhenglun Kong, Tianyun Zhang, Ji Li, Zhengang Li, Hang Liu, and Caiwen Ding · 2009
Earlier work this paper cites.
Spatten: Efficient sparse attention architecture with cascade token and head pruning, 2020a
Hanrui Wang, Zhekai Zhang, and Song Han · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation, 2013
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Data-free parameter pruning for deep neural networks, 2015
Suraj Srinivas and R. Venkatesh Babu · 2015
Earlier work this paper cites.
Pruning filters for efficient convnets
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Learning structured sparsity in deep neural networks, 2016
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2016
Cited alongside, same era.
Designing energy-efficient convolutional neural networks using energy-aware pruning, 2016
Tien-Ju Yang, Yu-Hsin Chen, and Vivienne Sze · 2016
Cited alongside, same era.
Thinet: A filter level pruning method for deep neural network compression
Jian-Hao Luo, Jianxin Wu, and Weiyao Lin · 2017
Cited alongside, same era.
Nisp: Pruning networks using neuron importance score propagation, 2017
Ruichi Yu, Ang Li, Chun-Fu Chen, Jui-Hsin Lai, Vlad I. Morariu, Xintong Han, Mingfei Gao, Ching-Yung Lin, and Larry S. Davis · 2017
Cited alongside, same era.
Seq2sql: Generating structured queries from natural language using reinforcement learning
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021
Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman · 2021
Later among the works it cites.
Dsee: Dually sparsity-embedded efficient tuning of pre-trained language models, 2021
Xuxi Chen, Tianlong Chen, Yu Cheng, Weizhu Chen, Zhangyang Wang, and Ahmed Hassan Awadallah · 2021
Later among the works it cites.
Language model as an annotator: Exploring dialogpt for dialogue summarization, 2021
Xiachong Feng, Xiaocheng Feng, Libo Qin, Bing Qin, and Ting Liu · 2021
Later among the works it cites.
Block pruning for faster transformers, 2021
François Lagunas, Ella Charlaix, Victor Sanh, and Alexander M. Rush · 2021
Later among the works it cites.
A short study on compressing decoder-based language models, 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Victor Zhong, Caiming Xiong, and Richard Socher · 2017
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression, 2017
Michael Zhu and Suyog Gupta · 2017
Cited alongside, same era.
Regularizing deep neural networks by enhancing diversity in feature extraction
Babajide O. Ayinde, Tamer Inanc, and Jacek M. Zurada · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2018
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Cited alongside, same era.
Rethinking the value of network pruning, 2018
Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell · 2018
Cited alongside, same era.
Piggyback: Adapting a single network to multiple tasks by learning to mask weights, 2018
Arun Mallya, Dillon Davis, and Svetlana Lazebnik · 2018
Cited alongside, same era.
Improving deep neural network sparsity through decorrelation regularization
Xiaotian Zhu, Wengang Zhou, and Houqiang Li · 2018
Cited alongside, same era.
Tianda Li, Yassir El Mesbahi, Ivan Kobyzev, Ahmad Rashid, Atif Mahmud, Nithin Anchuri, Habib Hajimolahoseini, Yang Liu, and Mehdi Rezagholizadeh · 2021
Later among the works it cites.
When to prune? a policy towards early structural pruning
Maying Shen, Pavlo Molchanov, Hongxu Yin, and Jose M Alvarez · 2021
Later among the works it cites.
Rethinking network pruning – under the pre-train and fine-tune paradigm, 2021
Dongkuan Xu, Ian E. H. Yen, Jinxi Zhao, and Zhibin Xiao · 2021
Later among the works it cites.
Leap: Learnable pruning for transformer-based models, 2021
Zhewei Yao, Xiaoxia Wu, Linjian Ma, Sheng Shen, Kurt Keutzer, Michael W. Mahoney, and Yuxiong He · 2021
Later among the works it cites.
Prune once for all: Sparse pre-trained language models, 2021
Ofir Zafrir, Ariel Larey, Guy Boudoukh, Haihao Shen, and Moshe Wasserblat · 2021
Later among the works it cites.
Kronecker decomposition for GPT compression
Ali Edalati, Marzieh Tahaei, Ahmad Rashid, Vahid Nia, James Clark, and Mehdi Rezagholizadeh · 2022
Later among the works it cites.
A fast post-training pruning framework for transformers, 2022
Woosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami · 2022
Later among the works it cites.
Parameter-efficient sparsity for large language models fine-tuning, 2022
Yuchao Li, Fuli Luo, Chuanqi Tan, Mengdi Wang, Songfang Huang, Shen Li, and Junjie Bai · 2022
Later among the works it cites.
Locating and editing factual associations in gpt, 2022
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Later among the works it cites.
Compression of generative pre-trained language models via quantization
Chaofan Tao, Lu Hou, Wei Zhang, Lifeng Shang, Xin Jiang, Qun Liu, Ping Luo, and Ngai Wong · 2022
Later among the works it cites.
Structured pruning learns compact and accurate models, 2022
Mengzhou Xia, Zexuan Zhong, and Danqi Chen · 2022
Later among the works it cites.
Can model compression improve nlp fairness, 2022
Guangxuan Xu and Qingyuan Hu · 2022
Later among the works it cites.
Platon: Pruning large transformer models with upper confidence bound of weight importance, 2022
Qingru Zhang, Simiao Zuo, Chen Liang, Alexander Bukharin, Pengcheng He, Weizhu Chen, and Tuo Zhao · 2022
Later among the works it cites.