Fetching the paper…
Reading the bibliography…
In recent years, the growing size of neural networks has led to a vast amount of research concerning compression techniques to mitigate the drawbacks of such large sizes.
Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
Tang, R.; Lu, Y.; Liu, L.; Mou, L.; Vechtomova, O.; and Lin, J. 2019 · 1903
Earlier work this paper cites.
Zmora, N.; Jacob, G.; Zlotnik, L.; Elharar, B.; and Novik, G. 2019 · 1910
Earlier work this paper cites.
Optimal brain damage
LeCun, Y.; Denker, J. S.; ; and Solla, S. A. 1990 · 1990
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
Deep Model Compression: Distilling Knowledge from Noisy Teachers
Sau, B. B.; and Balasubramanian, V. N. 2016 · 2016
Earlier work this paper cites.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Frankle, J.; and Carbin, M. 2018 · 2018
Cited alongside, same era.
Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
Jacob, B.; Kligys, S.; Chen, B.; Zhu, M.; Tang, M.; Howard, A.; Adam, H.; and Kalenichenko, D. 2018 · 2018
Cited alongside, same era.
Deep Mutual Learning
Zhang, Y.; Xiang, T.; Hospedales, T. M.; and Lu, H. 2018 · 2018
Cited alongside, same era.
Improving Generalization and Robustness with Noisy Collaboration in Knowledge Distillation
Arani, E.; Sarfraz, F.; and Zonooz, B. 2019 · 2019
Cited alongside, same era.
Knowledge Distillation via Route Constrained Optimization
Jin, X.; Peng, B.; Wu, Y.; Liu, Y.; Liu, J.; Liang, D.; Yan, J.; and Hu, X. 2019 · 2019
Later among the works it cites.
Improved Knowledge Distillation via Teacher Assistant
Mirzadeh, S.-I.; Farajtabar, M.; Li, A.; Levine, N.; Matsukawa, A.; and Ghasemzadeh, H. 2019 · 2019
Later among the works it cites.
Preparing Lessons: Improve Knowledge Distillation with Better Supervision
Wen, T.; Lai, S.; and Qian, X. 2019 · 2019
Later among the works it cites.
Regularizing Class-Wise Predictions via Self-Knowledge Distillation
Yun, S.; Park, J.; Lee, K.; and Shin, J. 2020 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…