Fetching the paper…
Reading the bibliography…
With the success of deep neural networks, knowledge distillation which guides the learning of a small student network from a large teacher network is being actively studied for model compression and transfer learning.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Distillation the knowledge in a neural network
G. Hinton, O. Yinyals, and J. Dean · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
A. Romero, N. Ballas, S. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
S. Han, H. Mao, and W. J. Dally · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
D. Pathak, P. Kahenbuhl, J. Donahue, T. Darrell, and A. Efros · 2016
Earlier work this paper cites.
Quantized convolutional neural networks for mobile devices
J. Wu, C. Leng, Y. Wang, Q. Hu, and J. Cheng · 2016
Earlier work this paper cites.
Wide residual networks
S. Zagoruyko and N. Komodakis · 2016
Earlier work this paper cites.
Segnet: A deep convolutional encoder-decoder architecture for image segmentation
V. Badrinarayanan, A. Kendall, and R. Cipolla · 2017
Earlier work this paper cites.
Mask r-cnn
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Cited alongside, same era.
Densely connected convolutional networks
G. Huang, Z. Liu, L. van der Maaten, and K. Weinberger · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, and et. al · 2017
Cited alongside, same era.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
J. Yim, D. Joo, J. Bae, and J. Kim · 2017
Cited alongside, same era.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
S. Zagoruyko and N. Komodakis · 2017
Cited alongside, same era.
Knowledge distillation via route constrained optimization
X. Jin, B. Peng, Y. Wu, Y. Liu, J. Liu, D. Liang, J. Yan, and X. Hu · 2019
Later among the works it cites.
Relational knowledge distillation
W. Park, D. Kim, Y. Lu, and M. Cho · 2019
Later among the works it cites.
Correlation congruence for knowledge distillation
B. Peng, X. Jin, J. Liu, D. Li, Y. Wu, Y. Liu, S. Zhou, and Z. Zhang · 2019
Later among the works it cites.
Meal: Multi-model ensemble via adversarial learning
Z. Shen, Z. He, and X. Xue · 2019
Later among the works it cites.
Similarity-preserving knowledge distillation
F. Tung and G. Mori · 2019
Later among the works it cites.
Be your own teacher: Improve the performance of convolutional neural networks via self distillation
L. Zhang, J. Song, A. Gao, J. Chen, C. Bao, and K. Ma · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Furlanello, Z. C. Lipton, M. Tschannen, L. Itti, and A. Anandkumar · 2018
Cited alongside, same era.
Paraphrasing complex network: Network compression via factor transfer
J. Kim, S. Park, and N. Kwak · 2018
Cited alongside, same era.
Knowledge distillation by on-the-fly native ensemble
X. Lan, X. Zhu, and S. Gong · 2018
Cited alongside, same era.
Nisp: Pruning networks using neuron importance score propagation
R. Yu, A. Li, C. Chen, J. Lai, V. Morariu, X. Hand, M. Gao, C. Lin, and L. Davis · 2018
Cited alongside, same era.
Deep mutual learning
Y. Zhang, T. Xiang, T. Hospedales, and H. Lu · 2018
Cited alongside, same era.
Variational information distillation for knowledge transfer
S. Ahn, S. X., Hu, A. Damianous, N. D. Lawrence, and Z. Dai · 2019
Cited alongside, same era.
On the efficacy of knowledge distillation
J. Cho and B. Hariharan · 2019
Cited alongside, same era.
Densely distilled flow-based knowledge transfer in teacher-student framework for image classification
J.-H. Bae, D. Yeo, J. Yim, N.-S. Kim, C.-S. Pyo, and J. Kim · 2020
Closest in time.
Online knowledge distillation with diverse peers
D. Chen, J. Mei, C. Wang, Y. Feng, and C. Chen · 2020
Closest in time.
Online knowledge distillation via collaborative learning
Q. Guo, X. Wang, Y. Wu, Z. Yu, D. Liang, X. Hu, and P. Luo · 2020
Closest in time.
Ensemble distribution distillation
A. Malinin, B. Mlodozeniec, and M. Gales · 2020
Closest in time.
Improved knowledge distillation via teacher assistant
S. I. Mirzadeh, M. Farajtabar, A. Li, N. Levine, A. Matsukawa, and H. Ghasemzadeh · 2020
Closest in time.
Contrastive representation distillation
Y. Tian, D. Krishnan, and P. Isola · 2020
Closest in time.
Knowledge distillation meets self-supervision
G. Xu, Z. Liu, X. Li, and C. C. Loy · 2020
Closest in time.
Revisit knowledge distillation: a teacher-free framework
L. Yuan, F. E. Tay, G. Li, T. Wang, and J. Feng · 2020
Closest in time.