Fetching the paper…
Reading the bibliography…
Knowledge distillation (KD) has gained much attention due to its effectiveness in compressing large-scale pre-trained models.
Improved Knowledge Distillation via Teacher Assistant
Mirzadeh, S.-I.; Farajtabar, M.; Li, A.; Levine, N.; Matsukawa, A.; and Ghasemzadeh, H. 2019 · 1902
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
Pham, H.; Xie, Q.; Dai, Z.; and Le, Q. V. 2020 · 2003
Earlier work this paper cites.
Automatically Constructing a Corpus of Sentential Paraphrases
Dolan, W. B.; and Brockett, C. 2005 · 2005
Earlier work this paper cites.
Model compression
Bucila, C.; Caruana, R.; and Niculescu-Mizil, A. 2006 · 2006
Earlier work this paper cites.
Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across Domains
Pan, H.; Wang, C.; Qiu, M.; Zhang, Y.; Li, Y.; and Huang, J. 2020 · 2012
Earlier work this paper cites.
Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C. D.; Ng, A. Y.; and Potts, C. 2013 · 2013
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Hinton, G. E.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100, 000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Earlier work this paper cites.
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
Finn, C.; Abbeel, P.; and Levine, S. 2017 · 2017
Cited alongside, same era.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
On First-Order Meta-Learning Algorithms
Nichol, A.; Achiam, J.; and Schulman, J. 2018 · 2018
Cited alongside, same era.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, A.; Nangia, N.; and Bowman, S. R. 2018 · 2018
Cited alongside, same era.
SciBERT: A Pretrained Language Model for Scientific Text
Beltagy, I.; Lo, K.; and Cohan, A. 2019 · 2019
Cited alongside, same era.
Structural Scaffolds for Citation Intent Classification in Scientific Publications
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2019 · 2019
Later among the works it cites.
Neural Network Acceptability Judgments
Warstadt, A.; Singh, A.; and Bowman, S. R. 2019 · 2019
Later among the works it cites.
Label-Similarity Curriculum Learning
Dogan, Ü.; Deshmukh, A. A.; Machura, M.; and Igel, C. 2020 · 2020
Later among the works it cites.
TinyBERT: Distilling BERT for Natural Language Understanding
Jiao, X.; Yin, Y.; Shang, L.; Jiang, X.; Chen, X.; Li, L.; Wang, F.; and Liu, Q. 2020 · 2020
Later among the works it cites.
MetaDistiller: Network Self-Boosting via Meta-Learned Top-Down Distillation
Liu, B.; Rao, Y.; Lu, J.; Zhou, J.; and Hsieh, C. 2020 · 2020
Later among the works it cites.
BioBERTpt - A Portuguese Neural Language Model for Clinical Named Entity Recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cohan, A.; Ammar, W.; van Zuylen, M.; and Cady, F. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; Desmaison, A.; Köpf, A.; Yang, E.; DeVito, Z.; Raison, M.; Tejani, A.; Chilamkurthy, S.; Steiner, B.; Fang, L.; Bai, J.; and Chintala, S. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Cited alongside, same era.
Patient Knowledge Distillation for BERT Model Compression
Sun, S.; Cheng, Y.; Gan, Z.; and Liu, J. 2019 · 2019
Cited alongside, same era.
Schneider, E. T. R.; de Souza, J. V. A.; Knafou, J.; Oliveira, L. E. S. e.; Copara, J.; Gumiel, Y. B.; Oliveira, L. F. A. d.; Paraiso, E. C.; Teodoro, D.; and Barra, C. M. C. M. 2020 · 2020
Later among the works it cites.
MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
Sun, Z.; Yu, H.; Song, X.; Liu, R.; Yang, Y.; and Zhou, D. 2020 · 2020
Later among the works it cites.
Knowledge Distillation: A Survey
Gou, J.; Yu, B.; Maybank, S. J.; and Tao, D. 2021 · 2021
Closest in time.
Annealing Knowledge Distillation
Jafari, A.; Rezagholizadeh, M.; Sharma, P.; and Ghodsi, A. 2021 · 2021
Closest in time.