Fetching the paper…
Reading the bibliography…
Knowledge distillation (KD) is commonly deemed as an effective model compression technique in which a compact model (student) is trained under the supervision of a larger pretrained model or an ensemble of models (teacher).
C. Buciluǎ, R. Caruana, and A. Niculescu-Mizil, “Model compression,” in Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining . ACM, 2006, pp. 535–541
2006
Earlier work this paper cites.
C.-F. Tsai and C. Hung, “Automatically annotating images with keywords: A review of image annotation systems,” Recent Patents on Computer Science , vol. 1, no. 1, pp. 55–68, 2008
2008
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
A. Krizhevsky, V. Nair, and G. Hinton, “Cifar-10 (canadian institute for advanced research),” URL http://www. cs. toronto. edu/kriz/cifar. html , vol. 8, 2010
2010
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
B. Mozafari, P. Sarkar, M. Franklin, M. Jordan, and S. Madden, “Scaling up crowd-sourcing to very large datasets: a case for active learning,” Proceedings of the VLDB Endowment , vol. 8, no. 2, pp. 125–136, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition. computer vision and pattern recognition (cvpr),” in 2016 IEEE Conference on , vol. 5, 2015, p. 6
2015
Cited alongside, same era.
2015
Cited alongside, same era.
2016
Cited alongside, same era.
S. Zagoruyko and N. Komodakis, “Wide residual networks,” arXiv preprint arXiv:1605.07146 , 2016
2016
Cited alongside, same era.
A. Voulodimos, N. Doulamis, A. Doulamis, and E. Protopapadakis, “Deep learning for computer vision: A brief review,” Computational intelligence and neuroscience , vol. 2018, 2018
2018
Later among the works it cites.
Y. Zhang, T. Xiang, T. M. Hospedales, and H. Lu, “Deep mutual learning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 4320–4328
2018
Later among the works it cites.
X. Lan, X. Zhu, and S. Gong, “Knowledge distillation by on-the-fly native ensemble,” in Proceedings of the 32nd International Conference on Neural Information Processing Systems . Curran Associates Inc., 2018, pp. 7528–7538
2018
Later among the works it cites.
G. Van Horn, O. Mac Aodha, Y. Song, Y. Cui, C. Sun, A. Shepard, H. Adam, P. Perona, and S. Belongie, “The inaturalist species classification and detection dataset,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 8769–8778
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Z. Zhou, P. Mertikopoulos, N. Bambos, S. Boyd, and P. W. Glynn, “Stochastic mirror descent in variationally coherent optimization problems,” in Advances in Neural Information Processing Systems , 2017, pp. 7040–7049
2017
Cited alongside, same era.
2017
Cited alongside, same era.
D. Arpit, S. Jastrzębski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, T. Maharaj, A. Fischer, A. Courville, Y. Bengio et al. , “A closer look at memorization in deep networks,” in Proceedings of the 34th International Conference on Machine Learning-Volume 70 . JMLR. org, 2017, pp. 233–242
2017
Cited alongside, same era.
J. Yim, D. Joo, J. Bae, and J. Kim, “A gift from knowledge distillation: Fast optimization, network minimization and transfer learning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 4133–4141
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Q. Dong, S. Gong, and X. Zhu, “Imbalanced deep learning by minority class incremental rectification,” IEEE transactions on pattern analysis and machine intelligence , vol. 41, no. 6, pp. 1367–1381, 2018
2018
Later among the works it cites.
2019
Later among the works it cites.
J. M. Johnson and T. M. Khoshgoftaar, “Survey on deep learning with class imbalance,” Journal of Big Data , vol. 6, no. 1, p. 27, 2019
2019
Later among the works it cites.
B. Heo, M. Lee, S. Yun, and J. Y. Choi, “Knowledge distillation with adversarial samples supporting decision boundary,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, 2019, pp. 3771–3778
2019
Later among the works it cites.
W. Park, D. Kim, Y. Lu, and M. Cho, “Relational knowledge distillation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 3967–3976
2019
Later among the works it cites.
2019
Later among the works it cites.