Fetching the paper…
Reading the bibliography…
Knowledge Distillation (KD) aims to distill the knowledge of a cumbersome teacher model into a lightweight student model.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
Do deep nets really need to be deep?
J. Ba and R. Caruana · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Cited alongside, same era.
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Cited alongside, same era.
Regularizing neural networks by penalizing confident output distributions
G. Pereyra, G. Tucker, J. Chorowski, Ł. Kaiser, and G. Hinton · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Cited alongside, same era.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
J. Yim, D. Joo, J. Bae, and J. Kim · 2017
Cited alongside, same era.
T. Furlanello, Z. C. Lipton, M. Tschannen, L. Itti, and A. Anandkumar · 2018
Later among the works it cites.
Shufflenet v2: Practical guidelines for efficient cnn architecture design
N. Ma, X. Zhang, H.-T. Zheng, and J. Sun · 2018
Later among the works it cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen · 2018
Later among the works it cites.
Deep mutual learning
Y. Zhang, T. Xiang, T. M. Hospedales, and H. Lu · 2018
Later among the works it cites.
Improved knowledge distillation via teacher assistant: Bridging the gap between student and teacher
S.-I. Mirzadeh, M. Farajtabar, A. Li, and H. Ghasemzadeh · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Visual relationship detection with internal and external linguistic knowledge distillation
R. Yu, A. Li, V. I. Morariu, and L. S. Davis · 2017
Cited alongside, same era.
Large scale distributed neural network training through online distillation
R. Anil, G. Pereyra, A. Passos, R. Ormandi, G. E. Dahl, and G. E. Hinton · 2018
Cited alongside, same era.
R. Müller, S. Kornblith, and G. Hinton · 2019
Closest in time.
Distilling object detectors with fine-grained feature imitation
T. Wang, L. Yuan, X. Zhang, and J. Feng · 2019
Closest in time.