Fetching the paper…
Reading the bibliography…
Knowledge Distillation (KD) transfers the knowledge from a high-capacity teacher network to strengthen a smaller student.
Har-Peled, S., Kushal, A.: Smaller coresets for k-median and k-means clustering. Discrete & Computational Geometry 37
2007
Earlier work this paper cites.
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 248–255 (2009)
2009
Earlier work this paper cites.
Krizhevsky, A., Hinton, G., et al.: Learning multiple layers of features from tiny images (2009)
2009
Earlier work this paper cites.
Olvera-López, J.A., Carrasco-Ochoa, J.A., Martínez-Trinidad, J.F., Kittler, J.: A review of instance selection methods. Artificial Intelligence Review 34
2010
Earlier work this paper cites.
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
Zagoruyko, S., Komodakis, N.: Wide residual networks. arXiv preprint arXiv:1605.07146 (2016)
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
Komodakis, N., Zagoruyko, S.: Paying more attention to attention: improving the performance of convolutional neural networks via attention transfer. In: Proceedings of the International Conference of Learning Representation (ICLR) (2017)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Han, B., Yao, Q., Yu, X., Niu, G., Xu, M., Hu, W., Tsang, I., Sugiyama, M.: Co-teaching: Robust training of deep neural networks with extremely noisy labels. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) 31
2018
Earlier work this paper cites.
Katharopoulos, A., Fleuret, F.: Not all samples are created equal: Deep learning with importance sampling. In: Proceedings of the International Conference on Machine Learning (ICML). pp. 2525–2534 (2018)
2018
Cited alongside, same era.
Ma, N., Zhang, X., Zheng, H.T., Sun, J.: Shufflenet v2: Practical guidelines for efficient cnn architecture design. In: Proceedings of the European Conference on Computer Vision (ECCV). pp. 116–131 (2018)
2018
Cited alongside, same era.
Passalis, N., Tefas, A.: Learning deep representations with probabilistic knowledge transfer. In: Proceedings of the European Conference on Computer Vision (ECCV). pp. 268–284 (2018)
2018
Cited alongside, same era.
Toneva, M., Sordoni, A., des Combes, R.T., Trischler, A., Bengio, Y., Gordon, G.J.: An empirical study of example forgetting during deep neural network learning. In: Proceedings of the International Conference of Learning Representation (ICLR) (2018)
2018
Cited alongside, same era.
Lin, M., Ji, R., Wang, Y., Zhang, Y., Zhang, B., Tian, Y., Shao, L.: Hrank: Filter pruning using high-rank feature map. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 1529–1538 (2020)
2020
Later among the works it cites.
Mirzasoleiman, B., Bilmes, J., Leskovec, J.: Coresets for data-efficient training of machine learning models. In: Proceedings of the International Conference on Machine Learning (ICML). pp. 6950–6960 (2020)
2020
Later among the works it cites.
Shen, Z., Liu, Z., Xu, D., Chen, Z., Cheng, K.T., Savvides, M.: Is label smoothing truly incompatible with knowledge distillation: An empirical study. In: Proceedings of the International Conference of Learning Representation (ICLR) (2020)
2020
Later among the works it cites.
Xu, G., Liu, Z., Li, X., Loy, C.C.: Knowledge distillation meets self-supervision. In: Proceedings of the European Conference on Computer Vision (ECCV). pp. 588–604 (2020)
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
Zhang, X., Zhou, X., Lin, M., Sun, J.: Shufflenet: An extremely efficient convolutional neural network for mobile devices. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 6848–6856 (2018)
2018
Cited alongside, same era.
Zhu, X., Gong, S., et al.: Knowledge distillation by on-the-fly native ensemble. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) (2018)
2018
Cited alongside, same era.
Ahn, S., Hu, S.X., Damianou, A., Lawrence, N.D., Dai, Z.: Variational information distillation for knowledge transfer. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 9163–9171 (2019)
2019
Cited alongside, same era.
Cai, H., Zhu, L., Han, S.: Proxylessnas: Direct neural architecture search on target task and hardware. In: Proceedings of the International Conference of Learning Representation (ICLR) (2019)
2019
Cited alongside, same era.
Heo, B., Kim, J., Yun, S., Park, H., Kwak, N., Choi, J.Y.: A comprehensive overhaul of feature distillation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 1921–1930 (2019)
2019
Cited alongside, same era.
Müller, R., Kornblith, S., Hinton, G.E.: When does label smoothing help? Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) (2019)
2019
Cited alongside, same era.
Park, W., Kim, D., Lu, Y., Cho, M.: Relational knowledge distillation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 3967–3976 (2019)
2019
Cited alongside, same era.
2020
Later among the works it cites.
Yin, H., Molchanov, P., Alvarez, J.M., Li, Z., Mallya, A., Hoiem, D., Jha, N.K., Kautz, J.: Dreaming to distill: Data-free knowledge transfer via deepinversion. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2020)
2020
Later among the works it cites.
Zhao, B., Mopuri, K.R., Bilen, H.: Dataset condensation with gradient matching. In: Proceedings of the International Conference of Learning Representation (ICLR) (2020)
2020
Later among the works it cites.
Chen, L., Wang, D., Gan, Z., Liu, J., Henao, R., Carin, L.: Wasserstein contrastive representation distillation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 16296–16305 (2021)
2021
Later among the works it cites.
Chen, P., Liu, S., Zhao, H., Jia, J.: Distilling knowledge via knowledge review. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 5008–5017 (2021)
2021
Later among the works it cites.
Yamamoto, K.: Learnable companding quantization for accurate low-bit neural networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). pp. 5029–5038 (2021)
2021
Later among the works it cites.
Zhang, Z., Chen, X., Chen, T., Wang, Z.: Efficient lottery ticket finding: Less data is more. In: Proceedings of the International Conference on Machine Learning (ICML). pp. 12380–12390 (2021)
2021
Later among the works it cites.
Li, S., Lin, M., Wang, Y., Fei, C., Shao, L., Ji, R.: Learning efficient gans for image translation via differentiable masks and co-attention distillation. IEEE Transactions on Multimedia (TMM) (2022)
2022
Closest in time.
Li, S., Lin, M., Wang, Y., Wu, Y., Tian, Y., Shao, L., Ji, R.: Distilling a powerful student model via online knowledge distillation. IEEE Transactions on Neural Networks and Learning Systems (TNNLS) (2022)
2022
Closest in time.