Fetching the paper…
Reading the bibliography…
Transformer attracts much attention because of its ability to learn global relations and superior performance.
A comprehensive overhaul of feature distillation
Heo, B., Kim, J., Yun, S., Park, H., Kwak, N., Choi, J.Y.: · 1930
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G.: · 2009
Earlier work this paper cites.
Do deep nets really need to be deep?
Ba, L.J., Caruana, R.: · 2013
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S.E., Chassang, A., Gatta, C., Bengio, Y.: · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: · 2014
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A.C., Fei-Fei, L.: · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., Dean, J.: · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., Sun, J.: · 2015
Earlier work this paper cites.
Zagoruyko, S., Komodakis, N.: · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., Sun, J.: · 2016
Earlier work this paper cites.
A survey of model compression and acceleration for deep neural networks
Cheng, Y., Wang, D., Zhou, P., Zhang, T.: · 2017
Earlier work this paper cites.
Like what you like: Knowledge distill via neuron selectivity transfer
Huang, Z., Wang, N.: · 2017
Earlier work this paper cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Yim, J., Joo, D., Bae, J., Kim, J.: · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., Lerer, A.: · 2017
Cited alongside, same era.
Xception: Deep learning with depthwise separable convolutions
Chollet, F.: · 2017
Cited alongside, same era.
Mask r-cnn
He, K., Gkioxari, G., Dollár, P., Girshick, R.: · 2017
Cited alongside, same era.
A survey on deep transfer learning
Tan, C., Sun, F., Kong, T., Zhang, W., Yang, C., Liu, C.: · 2018
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., Chen, L.C.: · 2018
Cited alongside, same era.
Paraphrasing complex network: Network compression via factor transfer
Kim, J., Park, S., Kwak, N.: · 2018
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: · 2020
Later among the works it cites.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: · 2020
Later among the works it cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., Zhou, M.: · 2020
Later among the works it cites.
Knowledge distillation from internal representations
Aguilar, G., Ling, Y., Zhang, Y., Yao, B., Fan, X., Guo, C.: · 2020
Later among the works it cites.
Celeba-spoof: Large-scale face anti-spoofing dataset with rich annotations
Zhang, Y., Yin, Z., Li, Y., Yin, G., Yan, J., Shao, J., Liu, Z.: · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Knowledge distillation via instance relationship graph
Liu, Y., Cao, J., Li, B., Yuan, C., Hu, W., Li, Y., Duan, Y.: · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M., Le, Q.: · 2019
Cited alongside, same era.
Relational knowledge distillation
Park, W., Kim, D., Lu, Y., Cho, M.: · 2019
Cited alongside, same era.
Contrastive representation distillation
Tian, Y., Krishnan, D., Isola, P.: · 2019
Cited alongside, same era.
Pay attention to features, transfer learn faster cnns
Wang, K., Gao, X., Zhao, Y., Li, X., Dou, D., Xu, C.Z.: · 2019
Cited alongside, same era.
Knowledge transfer via distillation of activation boundaries formed by hidden neurons
Heo, B., Lee, M., Yun, S., Choi, J.Y.: · 2019
Cited alongside, same era.
Tree-like decision distillation
Song, J., Zhang, H., Wang, X., Xue, M., Chen, Y., Sun, L., Tao, D., Song, M.: · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., Jégou, H.: · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B.: · 2021
Later among the works it cites.
Distilling knowledge via knowledge review
Chen, P., Liu, S., Zhao, H., Jia, J.: · 2021
Later among the works it cites.
Lipschitz continuity guided knowledge distillation
Shang, Y., Duan, B., Zong, Z., Nie, L., Yan, Y.: · 2021
Later among the works it cites.
Tensorrt
Nvidia: · 2022
Closest in time.
Spot-adaptive knowledge distillation
Song, J., Chen, Y., Ye, J., Song, M.: · 2022
Closest in time.