Fetching the paper…
Reading the bibliography…
In Knowledge Distillation, the teacher is generally much larger than the student, making the solution of the teacher likely to be difficult for the student to learn.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Coates, A., Ng, A., and Lee, H · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y · 2011
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Learning face representation from scratch
Yi, D., Lei, Z., Liao, S., and Li, S. Z · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Unifying distillation and privileged information
Lopez-Paz, D., Bottou, L., Schölkopf, B., and Vapnik, V · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Naeini, M. P., Cooper, G., and Hauskrecht, M · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
The megaface benchmark: 1 million faces for recognition at scale
Kemelmacher-Shlizerman, I., Seitz, S. M., Miller, D., and Brossard, E · 2016
Earlier work this paper cites.
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Cited alongside, same era.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Zagoruyko, S. and Komodakis, N · 2017
Cited alongside, same era.
Mobilefacenets: Efficient cnns for accurate real-time face verification on mobile devices
Chen, S., Liu, Y., Gao, X., and Han, Z · 2018
Cited alongside, same era.
Born again neural networks
Furlanello, T., Lipton, Z. C., Tschannen, M., Itti, L., and Anandkumar, A · 2018
Cited alongside, same era.
Paraphrasing complex network: Network compression via factor transfer
Kim, J., Park, S., and Kwak, N · 2018
Cited alongside, same era.
Additive margin softmax for face verification
Wang, F., Cheng, J., Liu, W., and Liu, H · 2018
Be your own teacher: Improve the performance of convolutional neural networks via self distillation
Zhang, L., Song, J., Gao, A., Chen, J., Bao, C., and Ma, K · 2019
Later among the works it cites.
Online knowledge distillation with diverse peers
Chen, D., Mei, J.-P., Wang, C., Feng, Y., and Chen, C · 2020
Later among the works it cites.
Online knowledge distillation via collaborative learning
Guo, Q., Wang, X., Wu, Y., Yu, Z., Liang, D., Hu, X., and Luo, P · 2020
Later among the works it cites.
Why distillation helps: a statistical perspective
Menon, A. K., Rawat, A. S., Reddi, S. J., Kim, S., and Kumar, S · 2020
Later among the works it cites.
Improved knowledge distillation via teacher assistant
Mirzadeh, S. I., Farajtabar, M., Li, A., Levine, N., Matsukawa, A., and Ghasemzadeh, H · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep mutual learning
Zhang, Y., Xiang, T., Hospedales, T. M., and Lu, H · 2018
Cited alongside, same era.
Variational information distillation for knowledge transfer
Ahn, S., Hu, S. X., Damianou, A., Lawrence, N. D., and Dai, Z · 2019
Cited alongside, same era.
On the efficacy of knowledge distillation
Cho, J. H. and Hariharan, B · 2019
Cited alongside, same era.
Arcface: Additive angular margin loss for deep face recognition
Deng, J., Guo, J., Xue, N., and Zafeiriou, S · 2019
Cited alongside, same era.
Adaptive regularization of labels
Ding, Q., Wu, S., Sun, H., Guo, J., and Xia, S.-T · 2019
Cited alongside, same era.
Knowledge distillation via route constrained optimization
Jin, X., Peng, B., Wu, Y., Liu, Y., Liu, J., Liang, D., Yan, J., and Hu, X · 2019
Cited alongside, same era.
Tian, Y., Krishnan, D., and Isola, P · 2020
Later among the works it cites.
Knowledge distillation meets self-supervision
Xu, G., Liu, Z., Li, X., and Loy, C. C · 2020
Later among the works it cites.
Knowledge transfer via dense cross-layer mutual-distillation
Yao, A. and Sun, D · 2020
Later among the works it cites.
Amln: adversarial-based mutual learning network for online knowledge distillation
Zhang, X., Lu, S., Gong, H., Luo, Z., and Liu, M · 2020
Later among the works it cites.
Distilling knowledge via knowledge review
Chen, P., Liu, S., Zhao, H., and Jia, J · 2021
Later among the works it cites.
Does knowledge distillation really work?
Stanton, S., Izmailov, P., Kirichenko, P., Alemi, A. A., and Wilson, A. G · 2021
Later among the works it cites.
Student customized knowledge distillation: Bridging the gap between student and teacher
Zhu, Y. and Wang, Y · 2021
Later among the works it cites.
Decoupled knowledge distillation
Zhao, B., Cui, Q., Song, R., Qiu, Y., and Liang, J · 2022
Later among the works it cites.