Fetching the paper…
Reading the bibliography…
Logit based knowledge distillation gets less attention in recent years since feature based methods perform better in most cases.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Trevor Darrell Jonathan Long, Evan Shelhamer · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Andrew Zisserman Karen Simonyan · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Evan Shelhamer, Jonathan Long, and Trevor Darrell · 2016
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Cited alongside, same era.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Sergey Zagoruyko and Nikos Komodakis · 2017
Cited alongside, same era.
Squeeze-and-excitation networks
Jie Hu, Li Shen, and Gang Sun · 2018
Cited alongside, same era.
Shufflenet V2: Practical guidelines for efficient cnn architecture design
Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun · 2018
Cited alongside, same era.
MobilenetV2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Cited alongside, same era.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
Contrastive representation distillation
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2020
Later among the works it cites.
Distilling knowledge via knowledge review
Pengguang Chen, Shu Liu, Hengshuang Zhao, and Jiaya Jia · 2021
Later among the works it cites.
Distilling global and local logits with densely connected relations
Youmin Kim, Jinbae Park, YounHo Jang, Muhammad Ali, Tae-Hyun Oh, and Sung-Ho Bae · 2021
Later among the works it cites.
Distilling holistic knowledge with graph neural networks
Zhou Sheng, Wang Yucheng, Chen Defang, Chen Jiawei, Wang Xin, Wang Can, and Bu Jiajun · 2021
Later among the works it cites.
Decoupled knowledge distillation
Zhao Borui, Cui Quan, Song Renjie, Qiu Yiyu, and Jiajun Liang · 2022
Later among the works it cites.
Cross-image relational knowledge distillation for semantic segmentation
Yang Chuanguang, Zhou Helong, An Zhulin, Jiang Xue, Xu Yongjun, and Zhang Qian · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang Xiangyu, Zhou Xinyu, Lin Mengxiao, and Sun Jian · 2018
Cited alongside, same era.
On the efficacy of knowledge distillation
Jang Hyun Cho and Bharath Hariharan · 2019
Cited alongside, same era.
A comprehensive overhaul of feature distillation
Byeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park, Nojun Kwak, and Jin Young Choi · 2019
Cited alongside, same era.
Relational knowledge distillation
Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
Snapshot distillation: Teacher-student optimization in one generation
Chenglin Yang, Lingxi Xie, Chi Su, and Alan L Yuille · 2019
Cited alongside, same era.
Later among the works it cites.
Knowledge distillation with the reused teacher classifier
Chen Defang, Mei Jian-Ping, Zhang Hailin, Wang Can, Feng Yan, and Chen Chun · 2022
Later among the works it cites.
Knowledge distillation via the target-aware transformer
Lin Sihao, Xie Hongwei, Wang Bing, Yu Kaicheng, Chang Xiaojun, Liang Xiaodan, and Wang Gang · 2022
Later among the works it cites.
Masked generative distillation, 2022
Zhendong Yang, Zhe Li, Mingqi Shao, Dachuan Shi, Zehuan Yuan, and Chun Yuan · 2022
Later among the works it cites.
Respecting transfer gap in knowledge distillation
Niu Yulei, Chen Long, Zhou Chang, and Zhang Hanwang · 2022
Later among the works it cites.