Fetching the paper…
Reading the bibliography…
Unlike existing knowledge distillation methods focus on the baseline settings, where the teacher models and training strategies are not that strong and competing as state-of-the-art approaches, this paper presents a method dubbed DIST to distill better from a stronger teacher.
A new measure of rank correlation
M. G. Kendall · 1938
Earlier work this paper cites.
The concise encyclopedia of statistics
Y. Dodge · 2008
Earlier work this paper cites.
Torchvision the machine-vision package of torch
S. Marcel and Y. Rodriguez · 2010
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis · 2016
Earlier work this paper cites.
Rethinking atrous convolution for semantic image segmentation
L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Earlier work this paper cites.
Learning from multiple teacher networks
S. You, C. Xu, C. Xu, and D. Tao · 2017
Earlier work this paper cites.
Pyramid scene parsing network
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia · 2017
Earlier work this paper cites.
Encoder-decoder with atrous separable convolution for semantic image segmentation
L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam · 2018
Earlier work this paper cites.
Paraphrasing complex network: Network compression via factor transfer
J. Kim, S. Park, and N. Kwak · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2018
Cited alongside, same era.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
X. Zhang, X. Zhou, M. Lin, and J. Sun · 2018
Cited alongside, same era.
Variational information distillation for knowledge transfer
S. Ahn, S. X. Hu, A. Damianou, N. D. Lawrence, and Z. Dai · 2019
Cited alongside, same era.
Cascade r-cnn: High quality object detection and instance segmentation
Z. Cai and N. Vasconcelos · 2019
Cited alongside, same era.
On the efficacy of knowledge distillation
J. H. Cho and B. Hariharan · 2019
Cited alongside, same era.
A comprehensive overhaul of feature distillation
B. Heo, J. Kim, S. Yun, H. Park, N. Kwak, and J. Y. Choi · 2019
Cited alongside, same era.
Improved knowledge distillation via teacher assistant
S. I. Mirzadeh, M. Farajtabar, A. Li, N. Levine, A. Matsukawa, and H. Ghasemzadeh · 2020
Later among the works it cites.
Probabilistic knowledge transfer for lightweight deep representation learning
N. Passalis, M. Tzelepi, and A. Tefas · 2020
Later among the works it cites.
Is label smoothing truly incompatible with knowledge distillation: An empirical study
Z. Shen, Z. Liu, D. Xu, Z. Chen, K.-T. Cheng, and M. Savvides · 2020
Later among the works it cites.
Intra-class feature variation distillation for semantic segmentation
Y. Wang, W. Zhou, T. Jiang, X. Bai, and Y. Xu · 2020
Later among the works it cites.
Knowledge distillation via softmax regression representation learning
J. Yang, B. Martinez, A. Bulat, and G. Tzimiropoulos · 2020
Later among the works it cites.
Greedynas: Towards fast one-shot nas with greedy supernet
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Knowledge transfer via distillation of activation boundaries formed by hidden neurons
B. Heo, M. Lee, S. Yun, and J. Y. Choi · 2019
Cited alongside, same era.
Relational knowledge distillation
W. Park, D. Kim, Y. Lu, and M. Cho · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Cited alongside, same era.
Correlation congruence for knowledge distillation
B. Peng, X. Jin, J. Liu, D. Li, Y. Wu, Y. Liu, S. Zhou, and Z. Zhang · 2019
Cited alongside, same era.
Efficientnet: Rethinking model scaling for convolutional neural networks
M. Tan and Q. Le · 2019
Cited alongside, same era.
Contrastive representation distillation
Y. Tian, D. Krishnan, and P. Isola · 2019
Cited alongside, same era.
S. You, T. Huang, M. Yang, F. Wang, C. Qian, and C. Zhang · 2020
Later among the works it cites.
Improve object detection with feature-based knowledge distillation: Towards accurate and efficient detectors
L. Zhang and K. Ma · 2020
Later among the works it cites.
Distilling knowledge via knowledge review
P. Chen, S. Liu, H. Zhao, and J. Jia · 2021
Later among the works it cites.
Knowledge distillation: A survey
J. Gou, B. Yu, S. J. Maybank, and D. Tao · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
Channel-wise knowledge distillation for dense prediction
C. Shu, Y. Liu, J. Gao, Z. Yan, and C. Shen · 2021
Later among the works it cites.
Densely guided knowledge distillation using multiple teacher assistants
W. Son, J. Na, J. Choi, and W. Hwang · 2021
Later among the works it cites.
Resnet strikes back: An improved training procedure in timm
R. Wightman, H. Touvron, and H. Jégou · 2021
Later among the works it cites.
To smooth or not to smooth? on compatibility between label smoothing and knowledge distillation, 2022
K. Chandrasegaran, N.-T. Tran, Y. ZHAO, and N. man Cheung · 2022
Closest in time.
Relational surrogate loss learning
T. Huang, Z. Li, H. Lu, Y. Shan, S. Yang, Y. Feng, F. Wang, S. You, and C. Xu · 2022
Closest in time.
Dyrep: Bootstrapping training with dynamic re-parameterization
T. Huang, S. You, B. Zhang, Y. Du, F. Wang, C. Qian, and C. Xu · 2022
Closest in time.
Cross-image relational knowledge distillation for semantic segmentation
C. Yang, H. Zhou, Z. An, X. Jiang, Y. Xu, and Q. Zhang · 2022
Closest in time.