Fetching the paper…
Reading the bibliography…
In many practical scenarios -- like hyperparameter search or continual retraining with new data -- related training runs are performed many times in sequence.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., Dean, J., et al · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Learning from multiple teacher networks
You, S., Xu, C., Xu, C., and Tao, D · 2017
Earlier work this paper cites.
Critical learning periods in deep networks
Achille, A., Rovere, M., and Soatto, S · 2019
Cited alongside, same era.
Ensemble knowledge distillation for learning improved and efficient networks
Asif, U., Tang, J., and Harrer, S · 2019
Cited alongside, same era.
A two-teacher framework for knowledge distillation
Chen, X., Su, J., and Zhang, J · 2019
Cited alongside, same era.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Allen-Zhu, Z. and Li, Y · 2020
Later among the works it cites.
On warm-starting neural network training
Ash, J. and Adams, R. P · 2020
Later among the works it cites.
Feature-level ensemble knowledge distillation for aggregating knowledge from multiple networks
Park, S. and Kwak, N · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…