Fetching the paper…
Reading the bibliography…
The proposed pruning strategy offers merits over weight-based pruning techniques: (1) it avoids irregular memory access since representations and matrices can be squeezed into their smaller but dense counterparts, leading to greater speedup; (2) in a manner of top-down pruning, the proposed method operates from a more global perspective based on training signals in the top layer, and prunes each layer by propagating the effect of global signals through layers, leading to better performances at the same sparsity level.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker. 2019 · 1902
Earlier work this paper cites.
Energy and policy considerations for deep learning in nlp
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019a · 1907
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin. 2019 · 1909
Earlier work this paper cites.
Reweighted proximal pruning for large-scale language representation
Fu-Ming Guo, Sijia Liu, Finlay S Mungall, Xue Lin, and Yanzhi Wang. 2019 · 1909
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 1909
Earlier work this paper cites.
Pruning a bert-based question answering model
J Scott McCarley. 2019 · 1910
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla. 1990 · 1990
Earlier work this paper cites.
A practical approach to feature selection
Kenji Kira and Larry A Rendell. 1992 · 1992
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David Stork. 1993 · 1993
Earlier work this paper cites.
Compressing large-scale transformer-based models: A case study on bert
Prakhar Ganesh, Yao Chen, Xin Lou, Mohammad Ali Khan, Yin Yang, Deming Chen, Marianne Winslett, Hassan Sajjad, and Preslav Nakov. 2020 · 2002
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020 · 2003
Earlier work this paper cites.
An introduction to variable and feature selection
Isabelle Guyon and André Elisseeff. 2003 · 2003
Earlier work this paper cites.
Sac: Accelerating and structuring self-attention via sparse adaptive connection
Xiaoya Li, Yuxian Meng, Qinghong Han, Fei Wu, and Jiwei Li. 2020 · 2003
Earlier work this paper cites.
Feature selection, ¡i¿l¡/i¿¡sub¿1¡/sub¿ vs. ¡i¿l¡/i¿¡sub¿2¡/sub¿ regularization, and rotational invariance
Andrew Y. Ng. 2004 · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
Feature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy
Hanchuan Peng, Fuhui Long, and Chris Ding. 2005 · 2005
Earlier work this paper cites.
The lottery ticket hypothesis for pre-trained bert networks
Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, and Michael Carbin. 2020 · 2007
Earlier work this paper cites.
A stability index for feature selection
Ludmila I Kuncheva. 2007 · 2007
Earlier work this paper cites.
Stable feature selection via dense feature groups
Lei Yu, Chris Ding, and Steven Loscalzo. 2008 · 2008
Earlier work this paper cites.
Normalized mutual information feature selection
Pablo A Estévez, Michel Tesmer, Claudio A Perez, and Jacek M Zurada. 2009 · 2009
Earlier work this paper cites.
Feature selection with dynamic mutual information
Huawen Liu, Jigui Sun, Lei Liu, and Huijie Zhang. 2009 · 2009
Earlier work this paper cites.
Nearest neighbor machine translation
Urvashi Khandelwal, Angela Fan, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020 · 2010
Earlier work this paper cites.
High-dimensional ising model selection using l l 1 -regularized logistic regression
Pradeep Ravikumar, Martin J. Wainwright, and John D. Lafferty. 2010 · 2010
Earlier work this paper cites.
Conditional likelihood maximisation: a unifying framework for information theoretic feature selection
Gavin Brown, Adam Pocock, Ming-Jie Zhao, and Mikel Luján. 2012 · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
A survey on feature selection methods
Girish Chandrashekar and Ferat Sahin. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Learning sentiment-specific word embedding for twitter sentiment classification
Duyu Tang, Furu Wei, Nan Yang, Ming Zhou, Ting Liu, and Bing Qin. 2014 · 2014
Earlier work this paper cites.
A review of feature selection methods based on mutual information
Jorge R Vergara and Pablo A Estévez. 2014 · 2014
Cited alongside, same era.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015 · 2015
Cited alongside, same era.
Feature selection using joint mutual information maximisation
Mohamed Bennasar, Yulia Hicks, and Rossitza Setchi. 2015 · 2015
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015 · 2015
Cited alongside, same era.
Variational dropout and the local reparameterization trick
Durk P Kingma, Tim Salimans, and Max Welling. 2015 · 2015
Cited alongside, same era.
Learning the number of neurons in deep networks
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Later among the works it cites.
Data-driven sparse structure selection for deep neural networks
Zehao Huang and Naiyan Wang. 2018 · 2018
Later among the works it cites.
Snip: Single-shot network pruning based on connection sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip HS Torr. 2018 · 2018
Later among the works it cites.
To prune, or not to prune: Exploring the efficacy of pruning for model compression
Suyog Gupta Michael H. Zhu. 2018 · 2018
Later among the works it cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jose M Alvarez and Mathieu Salzmann. 2016 · 2016
Cited alongside, same era.
Feature selection for high-dimensional data
Verónica Bolón-Canedo, Noelia Sánchez-Maroño, and Amparo Alonso-Betanzos. 2016 · 2016
Cited alongside, same era.
Dynamic network surgery for efficient dnns
Yiwen Guo, Anbang Yao, and Yurong Chen. 2016 · 2016
Cited alongside, same era.
Network trimming: A data-driven neuron pruning approach towards efficient deep architectures
Hengyuan Hu, Rui Peng, Yu-Wing Tai, and Chi-Keung Tang. 2016 · 2016
Cited alongside, same era.
Fasttext. zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. 2016 · 2016
Cited alongside, same era.
Pruning filters for efficient convnets
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf. 2016 · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Feature selection via mutual information: New theoretical insights
Mario Beraha, Alberto Maria Metelli, Matteo Papini, Andrea Tirinzoni, and Marcello Restelli. 2019 · 2019
Later among the works it cites.
Approximated oracle filter pruning for destructive CNN width optimization
Xiaohan Ding, Guiguang Ding, Yuchen Guo, Jungong Han, and Chenggang Yan. 2019 · 2019
Later among the works it cites.
Network pruning via transformable architecture search
Xuanyi Dong and Yi Yang. 2019 · 2019
Later among the works it cites.
Learning sparse networks using targeted dropout
Aidan N. Gomez, Ivan Zhang, Siddhartha Rao Kamalakara, Divyam Madaan, Kevin Swersky, Yarin Gal, and Geoffrey E. Hinton. 2019 · 2019
Later among the works it cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Later among the works it cites.
Importance estimation for neural network pruning
Pavlo Molchanov, Arun Mallya, Stephen Tyree, Iuri Frosio, and Jan Kautz. 2019 · 2019
Later among the works it cites.
On variational bounds of mutual information
Ben Poole, Sherjil Ozair, Aaron Van Den Oord, Alex Alemi, and George Tucker. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Later among the works it cites.
Autoslim: Towards one-shot architecture search for channel numbers
Jiahui Yu and Thomas Huang. 2019 · 2019
Later among the works it cites.
Description based text classification with reinforcement learning
Duo Chai, Wei Wu, Qinghong Han, Fei Wu, and Jiwei Li. 2020 · 2020
Later among the works it cites.
Club: A contrastive log-ratio upper bound of mutual information
Pengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu, Zhe Gan, and Lawrence Carin. 2020 · 2020
Later among the works it cites.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin. 2020 · 2020
Later among the works it cites.
Compressing BERT: Studying the effects of weight pruning on transfer learning
Mitchell Gordon, Kevin Duh, and Nicholas Andrews. 2020 · 2020
Later among the works it cites.
Revisiting self-training for neural sequence generation
Junxian He, Jiatao Gu, Jiajun Shen, and Marc’Aurelio Ranzato. 2020 · 2020
Later among the works it cites.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander M. Rush. 2020 · 2020
Later among the works it cites.
Neural semi-supervised learning for text classification under large-scale pretraining
Zijun Sun, Chun Fan, Xiaofei Sun, Yuxian Meng, Fei Wu, and Jiwei Li. 2020 · 2020
Later among the works it cites.
Structured pruning of large language models
Ziheng Wang, Jeremy Wohlwend, and Tao Lei. 2020 · 2020
Later among the works it cites.
Bertgcn: Transductive text classification by combining gcn and bert
Yuxiao Lin, Yuxian Meng, Xiaofei Sun, Qinghong Han, Kun Kuang, Jiwei Li, and Fei Wu. 2021 · 2021
Closest in time.
Fast nearest neighbor machine translation
Yuxian Meng, Xiaoya Li, Xiayu Zheng, Fei Wu, Xiaofei Sun, Tianwei Zhang, and Jiwei Li. 2021 · 2021
Closest in time.
Chinesebert: Chinese pretraining enhanced by glyph and pinyin information
Zijun Sun, Xiaoya Li, Xiaofei Sun, Yuxian Meng, Xiang Ao, Qing He, Fei Wu, and Jiwei Li. 2021 · 2021
Closest in time.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
Hanrui Wang, Zhekai Zhang, and Song Han. 2021 · 2021
Closest in time.
Adaptive nearest neighbor machine translation
Xin Zheng, Zhirui Zhang, Junliang Guo, Shujian Huang, Boxing Chen, Weihua Luo, and Jiajun Chen. 2021 · 2021
Closest in time.
Learning structured sparsity in deep neural networks
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. 2016 · 2082
Closest in time.