Fetching the paper…
Reading the bibliography…
Deep pre-trained Transformer models have achieved state-of-the-art results over a variety of natural language processing (NLP) tasks.
Distilling task-specific knowledge from BERT into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019 · 1903
Earlier work this paper cites.
A signal propagation perspective for pruning neural networks at initialization
Namhoon Lee, Thalaiyasingam Ajanthan, Stephen Gould, and Philip HS Torr. 2019a · 1906
Earlier work this paper cites.
Sparse networks from scratch: Faster training without losing performance
Tim Dettmers and Luke Zettlemoyer. 2019 · 1907
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 1908
Earlier work this paper cites.
TinyBERT: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Universal text representation from bert: An empirical study
Xiaofei Ma, Peng Xu, Zhiguo Wang, Ramesh Nallapati, and Bing Xiang. 2019 · 1910
Earlier work this paper cites.
Pruning a BERT-based question answering model
JS McCarley. 2019 · 1910
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Structured pruning of large language models
Ziheng Wang, Jeremy Wohlwend, and Tao Lei. 2019b · 1910
Earlier work this paper cites.
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019 · 1910
Earlier work this paper cites.
On information and sufficiency
Solomon Kullback and Richard A Leibler. 1951 · 1951
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S Denker, and Sara A Solla. 1990 · 1990
Earlier work this paper cites.
AdaBERT: Task-adaptive BERT compression with differentiable neural architecture search
Daoyuan Chen, Yaliang Li, Minghui Qiu, Zhen Wang, Bofang Li, Bolin Ding, Hongbo Deng, Jun Huang, Wei Lin, and Jingren Zhou. 2020a · 2001
Earlier work this paper cites.
Compressing BERT: Studying the effects of weight pruning on transfer learning
Mitchell A Gordon, Kevin Duh, and Nicholas Andrews. 2020 · 2002
Earlier work this paper cites.
Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joseph E Gonzalez. 2020 · 2002
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2005
Cited alongside, same era.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett. 2005 · 2005
Cited alongside, same era.
The PASCAL recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2006 · 2006
Cited alongside, same era.
Senteval: An evaluation toolkit for universal sentence representations
Alexis Conneau and Douwe Kiela. 2018 · 2018
Later among the works it cites.
Learning sparse neural networks through L0 regularization
Christos Louizos, Max Welling, and Diederik P. Kingma. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, and Michael Carbin. 2020b · 2007
Cited alongside, same era.
Semeval-2012 task 6: A pilot on semantic textual similarity
Eneko Agirre, Mona Diab, Daniel Cer, and Aitor Gonzalez-Agirre. 2012 · 2012
Cited alongside, same era.
Semeval-2013 shared task: Semantic textual similarity
Eneko Agirre, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. 2013 · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
Semeval-2014 task 10: Multilingual semantic textual similarity
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2014 · 2014
Cited alongside, same era.
Semeval-2014 task 1: Evaluation of compositional distributional semantic models on full sentences through semantic relatedness and textual entailment
Marco Marelli, Luisa Bentivogli, Marco Baroni, Raffaella Bernardi, Stefano Menini, and Roberto Zamparelli. 2014 · 2014
Cited alongside, same era.
Semeval-2015 task 2: Semantic textual similarity, english, spanish and pilot on interpretability
Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Inigo Lopez-Gazpio, Montse Maritxalar, Rada Mihalcea, et al. 2015 · 2015
Cited alongside, same era.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally. 2015 · 2015
Cited alongside, same era.
Later among the works it cites.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Later among the works it cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Later among the works it cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Later among the works it cites.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin. 2020 · 2020
Closest in time.
ALBERT: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Closest in time.
Poor man’s BERT: Smaller and faster transformer models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. 2020 · 2020
Closest in time.
MobileBERT: a compact task-agnostic bert for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2020
Closest in time.