Fetching the paper…
Reading the bibliography…
Knowledge distillation (KD) which transfers the knowledge from a large teacher model to a small student model, has been widely used to compress the BERT model recently.
Distilling task-specific knowledge from bert into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin. 2019 · 1903
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 1908
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Q-bert: Hessian based ultra low precision quantization of bert
S. Shen, Z. Dong, J. Ye, L. Ma, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer. 2019 · 1909
Earlier work this paper cites.
Pruning a bert-based question answering model
J. S. McCarley. 2019 · 1910
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 1910
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
O. Zafrir, G. Boudoukh, P. Izsak, and M. Wasserblat. 2019 · 1910
Earlier work this paper cites.
Coupling weight elimination with genetic algorithms to reduce network size and preserve generalization
George Bebis, Michael Georgiopoulos, and Takis Kasparis. 1997 · 1997
Earlier work this paper cites.
An introduction to genetic algorithms
Melanie Mitchell. 1998 · 1998
Earlier work this paper cites.
Evolving artificial neural networks
Xin Yao. 1999 · 1999
Earlier work this paper cites.
Compressing bert: Studying the effects of weight pruning on transfer learning
M. A. Gordon, K. Duh, and N. Andrews. 2020 · 2002
Earlier work this paper cites.
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2002
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020a · 2002
Cited alongside, same era.
Bert-of-theseus: Compressing bert by progressive module replacing
Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou. 2020 · 2002
Cited alongside, same era.
Dynabert: Dynamic bert with adaptive width and depth
L. Hou, L. Shang, X. Jiang, and Q. Liu. 2020 · 2004
Cited alongside, same era.
Evolving normalization-activation layers
Hanxiao Liu, Andrew Brock, Karen Simonyan, and Quoc V Le. 2020a · 2004
Cited alongside, same era.
Fastbert: a self-distilling bert with adaptive inference time
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. 2017 · 2017
Later among the works it cites.
Genetic cnn
Lingxi Xie and Alan Yuille. 2017 · 2017
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman. 2018 · 2018
Later among the works it cites.
Universal transformers
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser. 2019 · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
A tensorized transformer for language modeling
X. Ma, P. Zhang, S. Zhang, N. Duan, Y. Hou, M. Zhou, and D. Song. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Liu, P. Zhou, Z. Zhao, Z. Wang, H. Deng, and Q. Ju. 2020b · 2004
Cited alongside, same era.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Z. Sun, H. Yu, X. Song, R. Liu, Y. Yang, and D. Zhou. 2020 · 2004
Cited alongside, same era.
A dynamic chain-like agent genetic algorithm for global numerical optimization and feature selection
Xiao-Ping Zeng, Yong-Ming Li, and Jian Qin. 2009 · 2009
Cited alongside, same era.
Bert-emd: Many-to-many layer mapping for bert compression with earth mover’s distance
Jianquan Li, Xiaokang Liu, Honghong Zhao, Ruifeng Xu, Min Yang, and Yaohong Jin. 2020 · 2010
Cited alongside, same era.
A new local search based hybrid genetic algorithm for feature selection
Md. Monirul Kabir, Md. Shahjahan, and Kazuyuki Murase. 2011 · 2011
Cited alongside, same era.
Know what you don’t need: Single-shot meta-pruning for attention heads
Zhengyan Zhang, Fanchao Qi, Zhiyuan Liu, Qun Liu, and Maosong Sun. 2020 · 2011
Cited alongside, same era.
Extracting linguistic rules from data sets using fuzzy logic and genetic algorithms
Dan Meng and Zheng Pei. 2012 · 2012
Cited alongside, same era.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
The evolved transformer
David R So, Chen Liang, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Small and practical bert models for sequence labeling
Henry Tsai, Jason Riesa, Melvin Johnson, Naveen Arivazhagan, Xin Li, and Amelia Archer. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020 · 2020
Closest in time.
Depth-adaptive transformer
M. Elbayad, J. Gu, E. Grave, and M. Auli. 2020 · 2020
Closest in time.
Reducing transformer depth on demand with structured dropout
Angela F., Edouard G., and Armand J. 2020 · 2020
Closest in time.
Poor man’s bert: Smaller and faster transformer models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov. 2020 · 2020
Closest in time.