B. Heo, J. Kim, S. Yun, H. Park, N. Kwak, and J. Y. Choi, “A comprehensive overhaul of feature distillation,” in Proceedings of the IEEE International Conference on Computer Vision , 2019, pp. 1921–1930
1930
Earlier work this paper cites.
C. Bucilua, R. Caruana, and A. Niculescu-Mizil, “Model compression,” in Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining , 2006, pp. 535–541
2006
Earlier work this paper cites.
J. Ba and R. Caruana, “Do deep nets really need to be deep?” in Advances in neural information processing systems , 2014, pp. 2654–2662
2014
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” NIPS Workshop , 2014
2014
Earlier work this paper cites.
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio, “Fitnets: Hints for thin deep nets,” International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
Y. Le and X. Yang, “Tiny imagenet visual recognition challenge,” CS 231N , 2015
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al. , “Imagenet large scale visual recognition challenge,” International journal of computer vision , vol. 115, no. 3, pp. 211–252, 2015
2015
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 2818–2826
2016
Earlier work this paper cites.
L. Xie, J. Wang, Z. Wei, M. Wang, and Q. Tian, “Disturblabel: Regularizing cnn on the loss layer,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 4753–4762
2016
Earlier work this paper cites.
A. Shrivastava, A. Gupta, and R. Girshick, “Training region-based object detectors with online hard example mining,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 761–769
2016
Earlier work this paper cites.
B. B. Sau and V. N. Balasubramanian, “Deep model compression: Distilling knowledge from noisy teachers,” arXiv preprint arXiv:1610.09650 , 2016
Original
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in European conference on computer vision . Springer, 2016, pp. 630–645
2016
Earlier work this paper cites.
——, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Sgdr: Stochastic gradient descent with warm restarts,” arXiv preprint arXiv:1608.03983 , 2016
Original
2016
Earlier work this paper cites.
Z. Huang and N. Wang, “Like what you like: Knowledge distill via neuron selectivity transfer,” arXiv preprint arXiv:1707.01219 , 2017
Original
2017
Earlier work this paper cites.
J. Yim, D. Joo, J. Bae, and J. Kim, “A gift from knowledge distillation: Fast optimization, network minimization and transfer learning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 4133–4141
2017
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis, “Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer,” International Conference on Learning Representations , 2017
2017
Earlier work this paper cites.