Fetching the paper…
Reading the bibliography…
As language models increase in size by the day, methods for efficient inference are critical to leveraging their capabilities for various applications.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 1905
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 1908
Earlier work this paper cites.
Well-read students learn better: On the importance of pre-training compact models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 1908
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin. 2019 · 1909
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Structured pruning of a bert-based question answering model
JS McCarley, Rishav Chakravarti, and Avirup Sil. 2019 · 1910
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Structured pruning of large language models
Ziheng Wang, Jeremy Wohlwend, and Tao Lei. 2019 · 1910
Earlier work this paper cites.
Interpolation and approximation
Philip J Davis. 1975 · 1975
Earlier work this paper cites.
Cubic convolution interpolation for digital image processing
Robert Keys. 1981 · 1981
Earlier work this paper cites.
Comparing biases for minimal network construction with back-propagation
Stephen Jose Hanson and Lorien Y Pratt. 1989 · 1989
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S Denker, and Sara A Solla. 1989 · 1989
Earlier work this paper cites.
Kriging: a method of interpolation for geographical information systems
Margaret A Oliver and Richard Webster. 1990 · 1990
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
Babak Hassibi, David G Stork, and Gregory J Wolff. 1993 · 1993
Earlier work this paper cites.
Survey: interpolation methods in medical image processing
T.M. Lehmann, C. Gonner, and K. Spitzer. 1999 · 1999
Earlier work this paper cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2004
Earlier work this paper cites.
When bert plays the lottery, all tickets are winning
Sai Prasanna, Anna Rogers, and Anna Rumshisky. 2020 · 2005
Earlier work this paper cites.
Performance prediction for convolutional neural networks in edge devices
Halima Bouzidi, Hamza Ouarnoughi, Smail Niar, and Abdessamad Ait El Cadi. 2020 · 2010
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al. 2015 · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Cited alongside, same era.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and William J. Dally. 2016 · 2016
Cited alongside, same era.
Neuralpower: Predict and deploy energy-efficient convolutional neural networks
Ermao Cai, Da-Cheng Juan, Dimitrios Stamoulis, and Diana Marculescu. 2017 · 2017
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song. 2019 · 2019
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2019 · 2019
Later among the works it cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Later among the works it cites.
Dynabert: Dynamic bert with adaptive width and depth
Lu Hou, Zhiqi Huang, Lifeng Shang, Xin Jiang, Xiao Chen, and Qun Liu. 2020 · 2020
Later among the works it cites.
Overparameterized neural networks implement associative memory
Adityanarayanan Radhakrishnan, Mikhail Belkin, and Caroline Uhler. 2020 · 2020
Later among the works it cites.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander Rush. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploring the regularity of sparse structure in convolutional neural networks
Huizi Mao, Song Han, Jeff Pool, Wenshuo Li, Xingyu Liu, Yu Wang, and William J Dally. 2017 · 2017
Cited alongside, same era.
Block-sparse recurrent neural networks
Sharan Narang, Eric Undersander, and Gregory Diamos. 2017 · 2017
Cited alongside, same era.
Paleo: A performance model for deep neural networks
Hang Qi, Evan R Sparks, and Ameet Talwalkar. 2017 · 2017
Cited alongside, same era.
Learning intrinsic sparse structures within long short-term memory
Wei Wen, Yuxiong He, Samyam Rajbhandari, Minjia Zhang, Wenhan Wang, Fang Liu, Bin Hu, Yiran Chen, and Hai Li. 2017 · 2017
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman. 2017 · 2017
Cited alongside, same era.
Scalpel: Customizing dnn pruning to the underlying hardware parallelism
Jiecao Yu, Andrew Lukefahr, David Palframan, Ganesh Dasika, Reetuparna Das, and Scott Mahlke. 2017 · 2017
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Michael Zhu and Suyog Gupta. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020 · 2020
Later among the works it cites.
Training independent subnetworks for robust prediction
Marton Havasi, Rodolphe Jenatton, Stanislav Fort, Jeremiah Zhe Liu, Jasper Snoek, Balaji Lakshminarayanan, Andrew Mingbo Dai, and Dustin Tran. 2021 · 2021
Later among the works it cites.
Sparse progressive distillation: Resolving overfitting under pretrain-and-finetune paradigm
Shaoyi Huang, Dongkuan Xu, Ian EH Yen, Sung-en Chang, Bingbing Li, Shiyang Chen, Mimi Xie, Hang Liu, and Caiwen Ding. 2021 · 2021
Later among the works it cites.
Block pruning for faster transformers
François Lagunas, Ella Charlaix, Victor Sanh, and Alexander M Rush. 2021 · 2021
Later among the works it cites.
Mixmo: Mixing multiple inputs for multiple outputs via deep subnetworks
Alexandre Ramé, Rémy Sun, and Matthieu Cord. 2021 · 2021
Later among the works it cites.
Mlpruning: A multilevel structured pruning framework for transformer-based models
Zhewei Yao, Linjian Ma, Sheng Shen, Kurt Keutzer, and Michael W Mahoney. 2021 · 2021
Later among the works it cites.
Accelerating dnn training with structured data gradient pruning
Bradley McDanel, Helia Dinh, and John Magallanes. 2022 · 2022
Later among the works it cites.
DataMUX: Data multiplexing for neural networks
Vishvak Murahari, Carlos E Jimenez, Runzhe Yang, and Karthik R Narasimhan. 2022 · 2022
Later among the works it cites.
Structured pruning learns compact and accurate models
Mengzhou Xia, Zexuan Zhong, and Danqi Chen. 2022 · 2022
Later among the works it cites.
Mux-plms: Pre-training language models with data multiplexing
Vishvak Murahari, Ameet Deshpande, Carlos E Jimenez, Izhak Shafran, Mingqiu Wang, Yuan Cao, and Karthik Narasimhan. 2023 · 2023
Closest in time.
On the effect of dropping layers of pre-trained transformer models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. 2023 · 2023
Closest in time.