Fetching the paper…
Reading the bibliography…
Knowledge distillation (KD), as an efficient and effective model compression technique, has been receiving considerable attention in deep learning.
A. Krizhevsky, and G. Hinton, “Learning multiple layers of features from tiny images,” Technical Report. , 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and F.-F. Li, “ImageNet: A large-scale hierarchical image database,” IEEE Conf. Comput. Vis. Pattern Recog. , 2009
2009
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” Adv. Neural Inform. Process. Syst. , 2012
2012
Earlier work this paper cites.
S. Han, J. Pool, J. Tran, and W. J. Dally, “Learning both Weights and Connections for Efficient Neural Networks,” Adv. Neural Inform. Process. Syst. , 2015
2015
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the Knowledge in a Neural Network,” arXiv: 1503.02531 , 2015
2015
Earlier work this paper cites.
A. Romero, N. Ballas, and S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio, “FitNets: Hints for Thin Deep Nets,” Int. Conf. Learn. Represent. , 2015
2015
Earlier work this paper cites.
K. Simonyan, and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” Int. Conf. Pattern Recog. , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi, “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” Eur. Conf. Comput. Vis. , 2016
2016
Earlier work this paper cites.
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Binarized Neural Networks,” Adv. Neural Inform. Process. Syst. , 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” IEEE Conf. Comput. Vis. Pattern Recog. , 2016
2016
Earlier work this paper cites.
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning Filters for Efficient ConvNets,” Int. Conf. Learn. Represent. , 2017
2017
Earlier work this paper cites.
J. Yim, D. Joo, J. Bae, and J. Kim, “A gift from knowledge distillation: Fast optimization, network minimization and transfer learning,” IEEE Conf. Comput. Vis. Pattern Recog. , 2017
2017
Earlier work this paper cites.
S. Hou, X. Liu, and Z. Wang, “DualNet: Learn Complementary Features for Image Recognition,” Int. Conf. Comput. Vis. , 2017
2017
Earlier work this paper cites.
S. You, C. Xu, C. Xu, and D. Tao, “Learning from Multiple Teacher Networks,” ACM SIGKDD , 2017
2017
Earlier work this paper cites.
F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer, “SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size,” Int. Conf. Pattern Recog. , 2017
2017
Earlier work this paper cites.
Y. Zhang, T. Xiang, T. M. Hospedales, and H. Lu, “Deep Mutual Learning,” IEEE Conf. Comput. Vis. Pattern Recog. , 2018
2018
Earlier work this paper cites.
X. Lan, X. Zhu, and S. Gong, “Knowledge Distillation by On-the-Fly Native Ensemble,” IAdv. Neural Inform. Process. Syst. , 2018
2018
Cited alongside, same era.
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted Residuals and Linear Bottlenecks,” IEEE Conf. Comput. Vis. Pattern Recog. , 2018
2018
Cited alongside, same era.
N. Ma, X. Zhang, H.-T. Zheng, and J. Sun, “ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design,” Eur. Conf. Comput. Vis. , 2018
2018
Cited alongside, same era.
F. Tung, and G. Mori, “Similarity-preserving knowledge distillation,” Int. Conf. Comput. Vis. , 2019
2019
Cited alongside, same era.
C. Yang, L. Xie, C. Su, and A. L. Yuille, “OSnapshot distillation: Teacher-student optimization in one generation,” IEEE Conf. Comput. Vis. Pattern Recog. , 2019
2019
H. Chen, Y. Wang, C. Xu, C. Xu, and D. Tao, “Learning student networks via feature embedding,” IEEE Trans. Neur. Net. Lear. , 2020
2020
Later among the works it cites.
Y. Liu, C. Shu, J. Wang, and C. Shen, “Structured knowledge distillation for dense prediction,” IEEE Trans. Pattern Anal. Mach. Intell. , 2020
2020
Later among the works it cites.
K. Li, L. Yu, S. Wang, and P. A. Heng, “Towards Cross-Modality Medical Image Segmentation with Online Mutual Knowledge Distillation,” AAAI , 2020
2020
Later among the works it cites.
A. Yao, and D. Sun, “Knowledge Transfer via Dense Cross-Layer Mutual-Distillation,” Eur. Conf. Comput. Vis. , 2020
2020
Later among the works it cites.
Q. Guo, X. Wang, Y. Wu, Z. Yu, D. Liang, X. Hu, and P. Luo, “Online Knowledge Distillation via Collaborative Learning,” IEEE Conf. Comput. Vis. Pattern Recog. , 2020
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
M. Phuong, and C. H. Lampert, “Distillation-based training for multi-exit architectures,” Int. Conf. Comput. Vis. , 2019
2019
Cited alongside, same era.
C. Shen, M. Xue, X. Wang, J. Song, L. Sun, and M. Song, “Customizing student networks from heterogeneous teachers via adaptive knowledge amalgamation,” Int. Conf. Comput. Vis. , 2019
2019
Cited alongside, same era.
L. Yu, V. O. Yazici, X. Liu, J. Weijer, Y. Cheng, and A. Ramisa, “Learning metrics from teachers: Compact networks for image embedding,” IEEE Conf. Comput. Vis. Pattern Recog. , 2019
2019
Cited alongside, same era.
L. Zhang, J. Song, A. Gao, J. Chen, C. Bao, and K. Ma, “Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation,” Int. Conf. Comput. Vis. , 2019
2019
Cited alongside, same era.
W. Park, D. Kim, Y. Lu, and M. Cho, “Relational Knowledge Distillation,” IEEE Conf. Comput. Vis. Pattern Recog. , 2019
2019
Cited alongside, same era.
Y. Liu, J. Cao, B. Li, C. Yuan, W. Hu, Y. Li, and Y. Duan, “Knowledge Distillation via Instance Relationship Graph,” IEEE Conf. Comput. Vis. Pattern Recog. , 2019
2019
Cited alongside, same era.
B. Peng, X. Jin, J. Liu, D. Li, Y. Wu, Y. Liu, S. Zhou, and Z. Zhang, “Correlation Congruence for Knowledge Distillation,” Int. Conf. Comput. Vis. , 2019
2019
Cited alongside, same era.
Later among the works it cites.
D. Walawalkar, Z. Shen, and M. Savvides, “Online Ensemble Model Compression using Knowledge Distillation,” Eur. Conf. Comput. Vis. , 2020
2020
Later among the works it cites.
T. Li, J. Li, Z. Liu, and C. Zhang, “Few sample knowledge distillation for efficient network compression,” IEEE Conf. Comput. Vis. Pattern Recog. , 2020
2020
Later among the works it cites.
N. Passalis, M. Tzelepi, and A. Tefas, “Heterogeneous Knowledge Distillation using Information Flow Modeling,” IEEE Conf. Comput. Vis. Pattern Recog. , 2020
2020
Later among the works it cites.
K. Xu, L. Rui, Y. Li, and L. Gu, “Feature Normalized Knowledge Distillation for Image Classification,” Eur. Conf. Comput. Vis. , 2020
2020
Later among the works it cites.
Y. Guan, P. Zhao, B. Wang, Y. Zhang, C. Yao, K. Bian, and J. Tang, “Differentiable Feature Aggregation Search for Knowledge Distillation,” Eur. Conf. Comput. Vis. , 2020
2020
Later among the works it cites.
S.-I. Mirzadeh, M. Farajtabar, A. Li, and H. Ghasemzadeh, “Improved Knowledge Distillation via Teacher Assistant,” AAAI , 2020
2020
Later among the works it cites.
L. Yuan, F. E. Tay, G. Li, T. Wang, and J. Feng, “Revisiting Knowledge Distillation via Label Smoothing Regularization,” IEEE Conf. Comput. Vis. Pattern Recog. , 2020
2020
Later among the works it cites.
S. Yun, J. Park, K. Lee, and J. Shin, “Regularizing Class-wise Predictions via Self-knowledge Distillation,” IEEE Conf. Comput. Vis. Pattern Recog. , 2020
2020
Later among the works it cites.
G. Wu and S. Gong, “Peer Collaborative Learning for Online Knowledge Distillation,” IarXiv: 2006.04147 , 2020
2020
Later among the works it cites.
D. Chen, J.-P. Mei, C. Wang, Y. Feng, and C. Chen, “Online Knowledge Distillation with Diverse Peers,” AAAI , 2020
2020
Later among the works it cites.
X. Wu, R. He, Y. Hu, and Z. Sun, “Learning an Evolutionary Embedding via Massive Knowledge Distillation,” Int. J. Comput. Vis. , 2020, pp. 2089–2106
2089
Closest in time.