Fetching the paper…
Reading the bibliography…
Recent studies pointed out that knowledge distillation (KD) suffers from two degradation problems, the teacher-student gap and the incompatibility with strong data augmentations, making it not applicable to training state-of-the-art models, which are trained with advanced augmentations.
Model compression
Buciluǎ, C., Caruana, R., and Niculescu-Mizil, A · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Like what you like: Knowledge distill via neuron selectivity transfer
Huang, Z. and Wang, N · 2017
Earlier work this paper cites.
Meta-sgd: Learning to learn quickly for few-shot learning
Li, Z., Zhou, F., Chen, F., and Li, H · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Yim, J., Joo, D., Bae, J., and Kim, J · 2017
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2017
Cited alongside, same era.
Paraphrasing complex network: Network compression via factor transfer
Kim, J., Park, S., and Kwak, N · 2018
Cited alongside, same era.
Meta-gradient reinforcement learning
Xu, Z., van Hasselt, H., and Silver, D · 2018
Cited alongside, same era.
Variational information distillation for knowledge transfer
Ahn, S., Hu, S. X., Damianou, A., Lawrence, N. D., and Dai, Z · 2019
Cited alongside, same era.
An empirical analysis of the impact of data augmentation on knowledge distillation
Das, D., Massa, H., Kulkarni, A., and Rekatsinas, T · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Later among the works it cites.
Improved knowledge distillation via teacher assistant
Mirzadeh, S. I., Farajtabar, M., Li, A., Levine, N., Matsukawa, A., and Ghasemzadeh, H · 2020
Later among the works it cites.
Beit: Bert pre-training of image transformers
Bao, H., Dong, L., and Wei, F · 2021
Later among the works it cites.
Knowledge distillation: A good teacher is patient and consistent
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the efficacy of knowledge distillation
Cho, J. H. and Hariharan, B · 2019
Cited alongside, same era.
Adaptive gradient-based meta-learning methods
Khodak, M., Balcan, M.-F., and Talwalkar, A · 2019
Cited alongside, same era.
Lit: Learned intermediate representation training for model compression
Koratana, A., Kang, D., Bailis, P., and Zaharia, M · 2019
Cited alongside, same era.
When does label smoothing help?
Müller, R., Kornblith, S., and Hinton, G · 2019
Cited alongside, same era.
Relational knowledge distillation
Park, W., Kim, D., Lu, Y., and Cho, M · 2019
Cited alongside, same era.
Contrastive representation distillation
Tian, Y., Krishnan, D., and Isola, P · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y · 2019
Cited alongside, same era.
Beyer, L., Zhai, X., Royer, A., Markeeva, L., Anil, R., and Kolesnikov, A · 2021
Later among the works it cites.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Later among the works it cites.
An empirical study of training self-supervised vision transformers
Chen, X., Xie, S., and He, K · 2021
Later among the works it cites.
Isotonic data augmentation for knowledge distillation
Cui, W. and Yan, S · 2021
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2021
Later among the works it cites.
Efficient vision transformers via fine-grained manifold distillation
Jia, D., Han, K., Wang, Y., Tang, Y., Guo, J., Zhang, C., and Tao, D · 2021
Later among the works it cites.
Meta pseudo labels
Pham, H., Dai, Z., Xie, Q., and Le, Q. V · 2021
Later among the works it cites.
Is label smoothing truly incompatible with knowledge distillation: An empirical study
Shen, Z., Liu, Z., Xu, D., Chen, Z., Cheng, K.-T., and Savvides, M · 2021
Later among the works it cites.
How to train your vit? data, augmentation, and regularization in vision transformers
Steiner, A., Kolesnikov, A., Zhai, X., Wightman, R., Uszkoreit, J., and Beyer, L · 2021
Later among the works it cites.