Fetching the paper…
Reading the bibliography…
Knowledge distillation (KD) is one of the most potent ways for model compression.
Model compression
C. Buciluǎ, R. Caruana, and A. Niculescu-Mizil · 2006
Earlier work this paper cites.
Visualizing data using t-sne
L. v. d. Maaten and G. Hinton · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Do deep nets really need to be deep?
J. Ba and R. Caruana · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2015
Earlier work this paper cites.
Net2net: Accelerating learning via knowledge transfer
T. Chen, I. Goodfellow, and J. Shlens · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Like what you like: Knowledge distill via neuron selectivity transfer
Z. Huang and N. Wang · 2017
Cited alongside, same era.
Cascade residual learning: A two-stage convolutional neural network for stereo matching
J. Pang, W. Sun, J. S. Ren, C. Yang, and Q. Yan · 2017
Cited alongside, same era.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
J. Yim, D. Joo, J. Bae, and J. Kim · 2017
Cited alongside, same era.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising
K. Zhang, W. Zuo, Y. Chen, D. Meng, and L. Zhang · 2017
Later among the works it cites.
N2n learning: Network to network compression via policy gradient reinforcement learning
A. Ashok, N. Rhinehart, F. Beainy, and K. M. Kitani · 2018
Later among the works it cites.
Adversarial network compression
V. Belagiannis, A. Farshad, and F. Galasso · 2018
Later among the works it cites.
An embarrassingly simple approach for knowledge distillation
M. Gao, Y. Shen, Q. Li, J. Yan, L. Wan, D. Lin, C. C. Loy, and X. Tang · 2018
Later among the works it cites.
Self-supervised knowledge distillation using singular value decomposition
S. H. Lee, D. H. Kim, and B. C. Song · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Zagoruyko and N. Komodakis · 2017
Cited alongside, same era.
Accurate image super-resolution using very deep convolutional networks
J. Kim, J. Kwon Lee, and K. Mu Lee
Cited in the paper.
Multimodal residual learning for visual qa
J.-H. Kim, S.-W. Lee, D. Kwak, M.-O. Heo, J. Kim, J.-W. Ha, and B.-T. Zhang
Cited in the paper.
Progressive blockwise knowledge distillation for neural network acceleration
H. Wang, H. Zhao, X. Li, and X. Tan
Cited in the paper.
Esrgan: Enhanced super-resolution generative adversarial networks
X. Wang, K. Yu, S. Wu, J. Gu, Y. Liu, C. Dong, C. C. Loy, Y. Qiao, and X. Tang
Cited in the paper.
S.-I. Mirzadeh, M. Farajtabar, A. Li, and H. Ghasemzadeh · 2019
Later among the works it cites.