B. Heo, J. Kim, S. Yun, H. Park, N. Kwak, and J. Y. Choi, “A comprehensive overhaul of feature distillation,” in International Conference on Computer Vision , 2019, pp. 1921–1930
1930
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
J. Hoffman, E. Tzeng, T. Park, J.-Y. Zhu, P. Isola, K. Saenko, A. Efros, and T. Darrell, “CyCADA: Cycle-consistent adversarial domain adaptation,” in International Conference on Machine Learning , 2018, pp. 1989–1998
1998
Earlier work this paper cites.
C. Buciluǎ, R. Caruana, and A. Niculescu-Mizil, “Model compression,” in ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2006, pp. 535–541
2006
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems , vol. 25, 2012, pp. 1097–1105
2012
Earlier work this paper cites.
I. Kuzborskij and F. Orabona, “Stability and hypothesis transfer learning,” in International Conference on Machine Learning . PMLR, 2013, pp. 942–950
2013
Earlier work this paper cites.
M. Suk and B. Prabhakaran, “Real-time mobile facial expression recognition system-a case study,” in IEEE Conference on Computer Vision and Pattern Recognition Workshop , 2014, pp. 132–137
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in Neural Information Processing Systems , vol. 27, 2014
2014
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” arXiv preprint arXiv:1503.02531 , 2015
Original
2015
Earlier work this paper cites.
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio, “FitNets: Hints for thin deep nets,” in International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
A. Mordvintsev, C. Olah, and M. Tyka, “Inceptionism: Going deeper into neural networks,” 2015
2015
Earlier work this paper cites.
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2015, p. 3431–3440
2015
Earlier work this paper cites.
Y. Goldberg, “A primer on neural network models for natural language processing,” Journal of Artificial Intelligence Research , vol. 57, pp. 345–420, 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
M. Treml, J. Arjona-Medina, T. Unterthiner, R. Durgesh, F. Friedmann, P. Schuberth, A. Mayr, M. Heusel, M. Hofmarcher, M. Widrich et al. , “Speeding up semantic segmentation for autonomous driving,” in Advances in Neural Information Processing Systems , 2016
2016
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2016, pp. 3213–3223
2016
Earlier work this paper cites.
B. Chidlovskii, S. Clinchant, and G. Csurka, “Domain adaptation in the absence of source domain data,” in ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2016, pp. 451–460
2016
Earlier work this paper cites.
Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky, “Domain-adversarial training of neural networks,” The journal of machine learning research , vol. 17, no. 1, pp. 2096–2030, 2016
2016
Earlier work this paper cites.
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2017, pp. 1492–1500
2017
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 39, no. 6, p. 1137–1149, 2017
2017
Earlier work this paper cites.
L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam, “Rethinking atrous convolution for semantic image segmentation,” arXiv preprint arXiv:1706.05587 , 2017
Original
2017
Earlier work this paper cites.
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2017, pp. 4700–4708
2017
Earlier work this paper cites.
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning filters for efficient convnets,” in International Conference on Learning Representations , 2017
2017
Earlier work this paper cites.
R. G. Lopes, S. Fenu, and T. Starner, “Data-free knowledge distillation for deep neural networks,” in Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis, “Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer,” in International Conference on Learning Representations , 2017
2017
Earlier work this paper cites.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “Mixup: Beyond empirical risk minimization,” in International Conference on Learning Representations , 2017
2017
Earlier work this paper cites.
K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath, “Deep reinforcement learning: A brief survey,” IEEE Signal Processing Magazine , vol. 34, no. 6, pp. 26–38, 2017
2017
Earlier work this paper cites.
Y. Zhou, S.-M. Moosavi-Dezfooli, N.-M. Cheung, and P. Frossard, “Adaptive quantization for deep neural network,” in AAAI Conference on Artificial Intelligence , 2018
2018
Earlier work this paper cites.
A. Polino, R. Pascanu, and D. Alistarh, “Model compression via distillation and quantization,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
Z. Luo, J.-T. Hsieh, L. Jiang, J. C. Niebles, and L. Fei-Fei, “Graph distillation for action detection with privileged modalities,” in European Conference on Computer Vision , 2018, pp. 166–183
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805 , 2018
Original
2018
Earlier work this paper cites.
A. R. Nelakurthi, R. Maciejewski, and J. He, “Source free domain adaptation using an off-the-shelf classifier,” in IEEE International Conference on Big Data . IEEE, 2018, pp. 140–145
2018
Earlier work this paper cites.
M. Wang and W. Deng, “Deep visual domain adaptation: A survey,” Neurocomputing , vol. 312, pp. 135–153, 2018
2018
Earlier work this paper cites.
J. Tang and K. Wang, “Ranking distillation: Learning compact ranking models with high performance for recommender system,” in ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2018, pp. 2289–2298
2018
Earlier work this paper cites.
Y. Zhang, T. Xiang, T. M. Hospedales, and H. Lu, “Deep mutual learning,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2018, pp. 4320–4328
2018
Earlier work this paper cites.
Y. Zou, Z. Yu, B. Vijaya Kumar, and J. Wang, “Unsupervised domain adaptation for semantic segmentation via class-balanced self-training,” in European Conference on Computer Vision , 2018, pp. 289–305
2018
Earlier work this paper cites.
M. Caron, P. Bojanowski, A. Joulin, and M. Douze, “Deep clustering for unsupervised learning of visual features,” in European Conference on Computer Vision , 2018, pp. 132–149
2018
Earlier work this paper cites.
J. Kim, Y. Bhalgat, J. Lee, C. Patel, and N. Kwak, “QKD: Quantization-aware knowledge distillation,” arXiv preprint arXiv:1911.12491 , 2019
Original
2019
Earlier work this paper cites.
L. Chen, C. Yu, and L. Chen, “A new knowledge distillation for incremental object detection,” in International Joint Conference on Neural Networks . IEEE, 2019, pp. 1–7
2019
Earlier work this paper cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances in Neural Information Processing Systems , vol. 32, pp. 8026–8037, 2019
2019
Earlier work this paper cites.
Y. Wu, A. Kirillov, F. Massa, W.-Y. Lo, and R. Girshick, “Detectron2,” https://github.com/facebookresearch/detectron2 , 2019
2019
Earlier work this paper cites.
K. Chen, J. Wang, J. Pang, Y. Cao, Y. Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Xu, Z. Zhang, D. Cheng, C. Zhu, T. Cheng, Q. Zhao, B. Li, X. Lu, R. Zhu, Y. Wu, J. Dai, J. Wang, J. Shi, W. Ouyang, C. C. Loy, and D. Lin, “MMDetection: Open mmlab detection toolbox and benchmark,” arXiv preprint arXiv:1906.07155 , 2019
Original
2019
Earlier work this paper cites.
G. K. Nayak, K. R. Mopuri, V. Shaj, V. B. Radhakrishnan, and A. Chakraborty, “Zero-shot knowledge distillation in deep networks,” in International Conference on Machine Learning . PMLR, 2019, pp. 4743–4751
2019
Earlier work this paper cites.
H. Chen, Y. Wang, C. Xu, Z. Yang, C. Liu, B. Shi, C. Xu, C. Xu, and Q. Tian, “Data-free learning of student networks,” in International Conference on Computer Vision , 2019, pp. 3514–3522
2019
Earlier work this paper cites.
P. Micaelli and A. Storkey, “Zero-shot knowledge transfer via adversarial belief matching,” in Advances in Neural Information Processing Systems , 2019
2019
Earlier work this paper cites.
K. Bhardwaj, N. Suda, and R. Marculescu, “Dream distillation: A data-independent model compression framework,” in International Conference on Machine Learning , 2019
2019
Earlier work this paper cites.
W. M. Kouw and M. Loog, “A review of domain adaptation without target labels,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 43, no. 3, pp. 766–785, 2019
2019
Earlier work this paper cites.
Y. Liu, K. Chen, C. Liu, Z. Qin, Z. Luo, and J. Wang, “Structured knowledge distillation for semantic segmentation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 2604–2613
2019
Earlier work this paper cites.
F. Zhang, X. Zhu, and M. Ye, “Fast human pose estimation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 3517–3526
2019
Earlier work this paper cites.
J. Deng, Y. Pan, T. Yao, W. Zhou, H. Li, and T. Mei, “Relation distillation networks for video object detection,” in International Conference on Computer Vision , 2019, pp. 7023–7032
2019
Earlier work this paper cites.
C. Yang, L. Xie, C. Su, and A. L. Yuille, “Snapshot distillation: Teacher-student optimization in one generation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 2859–2868
2019
Earlier work this paper cites.
T. Wang, L. Yuan, X. Zhang, and J. Feng, “Distilling object detectors with fine-grained feature imitation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 4933–4942
2019
Earlier work this paper cites.
W. Park, D. Kim, Y. Lu, and M. Cho, “Relational knowledge distillation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 3967–3976
2019
Earlier work this paper cites.
Y. Liu, J. Cao, B. Li, C. Yuan, W. Hu, Y. Li, and Y. Duan, “Knowledge distillation via instance relationship graph,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 7096–7104
2019
Earlier work this paper cites.
Y. Hou, Z. Ma, C. Liu, and C. C. Loy, “Learning lightweight lane detection cnns by self attention distillation,” in International Conference on Computer Vision , 2019, pp. 1013–1021
2019
Earlier work this paper cites.
S. Park and N. Kwak, “FEED: Feature-level ensemble for knowledge distillation,” arXiv preprint arXiv:1909.10754 , 2019
Original
2019
Earlier work this paper cites.
Y.-H. Tsai, K. Sohn, S. Schulter, and M. Chandraker, “Domain adaptation for structured output via discriminative patch representations,” in International Conference on Computer Vision , 2019, pp. 1456–1465
2019
Earlier work this paper cites.
T.-H. Vu, H. Jain, M. Bucher, M. Cord, and P. Pérez, “Advent: Adversarial entropy minimization for domain adaptation in semantic segmentation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 2517–2526
2019
Earlier work this paper cites.
Y. Li, L. Yuan, and N. Vasconcelos, “Bidirectional learning for domain adaptation of semantic segmentation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 6936–6945
2019
Earlier work this paper cites.
L. Zhu, Z. Liu, and S. Han, “Deep leakage from gradients,” in Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
J. Yoo, M. Cho, T. Kim, and U. Kang, “Knowledge extraction with no observable data,” in Advances in Neural Information Processing Systems , 2019, pp. 2701–2710
2019
Earlier work this paper cites.
M. Nagel, M. v. Baalen, T. Blankevoort, and M. Welling, “Data-free quantization through weight equalization and bias correction,” in International Conference on Computer Vision , 2019, pp. 1325–1334
2019
Earlier work this paper cites.
X. Li, S. Zhang, B. Jiang, Y. Qi, M. C. Chuah, and N. Bi, “DAC: Data-free automatic acceleration of convolutional networks,” in IEEE Winter Conference on Applications of Computer Vision , 2019, pp. 1598–1606
2019
Earlier work this paper cites.
S. Sun, Y. Cheng, Z. Gan, and J. Liu, “Patient knowledge distillation for ber model compression,” in Conference on Empirical Methods in Natural Language Processing , 2019, pp. 4323–4332
2019
Earlier work this paper cites.
G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter, “Continual lifelong learning with neural networks: A review,” Neural Networks , vol. 113, pp. 54–71, 2019
2019
Earlier work this paper cites.
J. Liang, R. He, Z. Sun, and T. Tan, “Distant supervised centroid shift: A simple and efficient approach to visual domain adaptation,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 2975–2984
2019
Earlier work this paper cites.
K. Han, A. Vedaldi, and A. Zisserman, “Learning to discover novel visual categories via deep transfer clustering,” in International Conference on Computer Vision , 2019, pp. 8401–8409
2019
Earlier work this paper cites.
S. Shi, X. Wang, and H. Li, “PointRCNN: 3d object proposal generation and detection from point cloud,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 770–779
2019
Earlier work this paper cites.
X. Peng, Q. Bai, X. Xia, Z. Huang, K. Saenko, and B. Wang, “Moment matching for multi-source domain adaptation,” in International Conference on Computer Vision , 2019, pp. 1406–1415
2019
Earlier work this paper cites.
X. Yu, T. Liu, M. Gong, K. Zhang, K. Batmanghelich, and D. Tao, “Label-noise robust domain adaptation,” in International Conference on Machine Learning . PMLR, 2020, pp. 10 913–10 924
2019
Earlier work this paper cites.
H. Cha, J. Park, H. Kim, M. Bennis, and S.-L. Kim, “Proxy experience replay: Federated distillation for distributed reinforcement learning,” IEEE Intelligent Systems , vol. 35, no. 4, pp. 94–101, 2020
2020
Earlier work this paper cites.