Fetching the paper…
Reading the bibliography…
Pre-trained Language Models (PLMs) have been successful for a wide range of natural language processing (NLP) tasks.
Superglue: A stickier benchmark for general-purpose language understanding systems
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 1905
Earlier work this paper cites.
Model compression
C. Buciluǎ, R. Caruana, and A. Niculescu-Mizil · 2006
Earlier work this paper cites.
Distilling the knowledge in a neural network, 2015
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Earlier work this paper cites.
Born again neural networks, 2018
T. Furlanello, Z. C. Lipton, M. Tschannen, L. Itti, and A. Anandkumar · 2018
Earlier work this paper cites.
Black-box generation of adversarial text sequences to evade deep learning classifiers
J. Gao, J. Lanchantin, M. L. Soffa, and Y. Qi · 2018
Earlier work this paper cites.
Deep contextualized word representations
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford and K. Narasimhan · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Distilgpt2, 2019
HuggingFace · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Cited alongside, same era.
Zero: Memory optimization towards training A trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2019
Cited alongside, same era.
Patient knowledge distillation for bert model compression, 2019
S. Sun, Y. Cheng, Z. Gan, and J. Liu · 2019
Cited alongside, same era.
Language models are few-shot learners, 2020
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
TinyBERT: Distilling BERT for natural language understanding
X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu · 2020
Cited alongside, same era.
Bert-of-theseus: Compressing bert by progressive module replacing
C. Xu, W. Zhou, T. Ge, F. Wei, and M. Zhou · 2020
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding, 2020
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. Salakhutdinov, and Q. V. Le · 2020
Later among the works it cites.
Regularizing class-wise predictions via self-knowledge distillation
S. Yun, J. Park, K. Lee, and J. Shin · 2020
Later among the works it cites.
Compression of deep learning models for text: A survey, 2021
M. Gupta and P. Agrawal · 2021
Closest in time.
Rail-kd: Random intermediate layer mapping for knowledge distillation, 2021
M. A. Haidar, N. Anchuri, M. Rezagholizadeh, A. Ghaddar, P. Langlais, and P. Poupart · 2021
Closest in time.
Annealing knowledge distillation, 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ALP-KD: Attention-Based Layer Projection for Knowledge Distillation, 2020
P. Passban, Y. Wu, M. Rezagholizadeh, and Q. Liu · 2020
Cited alongside, same era.
Towards zero-shot knowledge distillation for natural language processing
A. Rashid, V. Lioutas, A. Ghaddar, and M. Rezagholizadeh · 2020
Cited alongside, same era.
DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters , page 3505–3506
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He · 2020
Cited alongside, same era.
A primer in BERTology: What we know about how BERT works
A. Rogers, O. Kovaleva, and A. Rumshisky · 2020
Cited alongside, same era.
Poor man’s BERT: smaller and faster transformer models
H. Sajjad, F. Dalvi, N. Durrani, and P. Nakov · 2020
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020
V. Sanh, L. Debut, J. Chaumond, and T. Wolf · 2020
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding, 2019b
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman
Cited in the paper.
A. Jafari, M. Rezagholizadeh, P. Sharma, and A. Ghodsi · 2021
Closest in time.
Not far away, not so close: Sample efficient nearest neighbour data augmentation via MiniMax
E. Kamalloo, M. Rezagholizadeh, P. Passban, and A. Ghodsi · 2021
Closest in time.
How to select one among all? an extensive empirical study towards the robustness of knowledge distillation in natural language understanding, 2021
T. Li, A. Rashid, A. Jafari, P. Sharma, A. Ghodsi, and M. Rezagholizadeh · 2021
Closest in time.
Zero-infinity: Breaking the GPU memory wall for extreme scale deep learning
S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith, and Y. He · 2021
Closest in time.
MATE-KD: Masked adversarial TExt, a companion to knowledge distillation
A. Rashid, V. Lioutas, and M. Rezagholizadeh · 2021
Closest in time.
Zero-offload: Democratizing billion-scale model training
J. Ren, S. Rajbhandari, R. Y. Aminabadi, O. Ruwase, S. Yang, M. Zhang, D. Li, and Y. He · 2021
Closest in time.