Fetching the paper…
Reading the bibliography…
Pre-training has improved model accuracy for both classification and generation tasks at the cost of introducing much larger and slower models.
Distilling task-specific knowledge from BERT into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019 · 1903
Earlier work this paper cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 1905
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 1905
Earlier work this paper cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 1908
Earlier work this paper cites.
Tinybert: Distilling BERT for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Structured pruning of a bert-based question answering model
J. S. McCarley. 2019 · 1910
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Structured pruning of large language models
Ziheng Wang, Jeremy Wohlwend, and Tao Lei. 2019 · 1910
Earlier work this paper cites.
What’s hidden in a randomly weighted neural network?
Vivek Ramanujan, Mitchell Wortsman, Aniruddha Kembhavi, Ali Farhadi, and Mohammad Rastegari. 2019 · 1911
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S. Denker, and Sara A. Solla. 1989 · 1989
Earlier work this paper cites.
Compressing BERT: studying the effects of weight pruning on transfer learning
Mitchell A. Gordon, Kevin Duh, and Nicholas Andrews. 2020 · 2002
Earlier work this paper cites.
Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joseph E. Gonzalez. 2020 · 2002
Earlier work this paper cites.
Ladabert: Lightweight adaptation of BERT through hybrid model compression
Yihuan Mao, Yujing Wang, Chufan Wu, Chen Zhang, Yang Wang, Yaming Yang, Quanlu Zhang, Yunhai Tong, and Jing Bai. 2020 · 2004
Earlier work this paper cites.
Poor man’s BERT: smaller and faster transformer models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. 2020 · 2004
Cited alongside, same era.
Mobilebert: a compact task-agnostic BERT for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2004
Cited alongside, same era.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander M. Rush. 2020 · 2005
Cited alongside, same era.
Pre-trained summarization distillation
Sam Shleifer and Alexander M. Rush. 2020 · 2010
Cited alongside, same era.
Compression of neural machine translation models via pruning
Abigail See, Minh-Thang Luong, and Christopher D. Manning. 2016 · 2016
Later among the works it cites.
Quora question pairs
Zihan Chen, Hongbo Zhang, Xiaoji Zhang, and Leqi Zhao. 2018 · 2018
Later among the works it cites.
The lottery ticket hypothesis: Training pruned neural networks
Jonathan Frankle and Michael Carbin. 2018 · 2018
Later among the works it cites.
Learning sparse neural networks through l_0 regularization
Christos Louizos, Max Welling, and Diederik P. Kingma. 2018 · 2018
Later among the works it cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Demi Guo, Alexander M. Rush, and Yoon Kim. 2020 · 2012
Cited alongside, same era.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville. 2013 · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
Learning both weights and connections for efficient neural networks
Song Han, Jeff Pool, John Tran, and William J. Dally. 2015 · 2015
Cited alongside, same era.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomás Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Cited alongside, same era.
Auto-sizing neural networks: With applications to n-gram language models
Kenton Murray and David Chiang. 2015 · 2015
Cited alongside, same era.
Fasttext.zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hervé Jégou, and Tomás Mikolov. 2016 · 2016
Cited alongside, same era.
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Patient knowledge distillation for BERT model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 2019
Later among the works it cites.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin. 2020 · 2020
Later among the works it cites.
Fastformers: Highly efficient transformer models for natural language understanding
Young Jin Kim and Hany Hassan Awadalla. 2020 · 2020
Later among the works it cites.
BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Turing-nlg: A 17-billion-parameter language model by microsoft
Corby Rosset. 2020 · 2020
Later among the works it cites.