Fetching the paper…
Reading the bibliography…
Although Deep neural networks (DNNs) have shown a strong capacity to solve large-scale problems in many areas, such DNNs are hard to be deployed in real-world systems due to their voluminous parameters.
K. Torkkola, “Feature extraction by non-parametric mutual information maximization,” JMLR , 2003
2003
Earlier work this paper cites.
D. Barber and F. Agakov, “The im algorithm: a variational approach to information maximization,” NeurIPS , 2004
2004
Earlier work this paper cites.
C. Buciluǎ, R. Caruana, and A. Niculescu-Mizil, “Model compression,” in KDD , 2006
2006
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in CVPR . IEEE, 2009
2009
Earlier work this paper cites.
A. Krizhevsky, G. Hinton et al. , “Learning multiple layers of features from tiny images,” 2009
2009
Earlier work this paper cites.
Y. Bengio, A. Courville, and P. Vincent, “Representation learning: A review and new perspectives,” TPAMI , 2013
2013
Earlier work this paper cites.
A. Karatzoglou, L. Baltrunas, and Y. Shi, “Learning to rank for recommender systems,” in RecSys . ACM, 2013
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” NeurIPS , 2014
2014
Earlier work this paper cites.
J. Li, L. Deng, Y. Gong, and R. Haeb-Umbach, “An overview of noise-robust automatic speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2014
2014
Earlier work this paper cites.
J. Ba and R. Caruana, “Deep model compression: compressing deep nets into 5% of their original size,” 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” NeurIPS , 2015
2015
Earlier work this paper cites.
T.-Y. Lin, A. RoyChowdhury, and S. Maji, “Bilinear cnn models for fine-grained visual recognition,” in ICCV , 2015
2015
Earlier work this paper cites.
A. Mahendran and A. Vedaldi, “Understanding deep image representations by inverting them,” in CVPR , 2015
2015
Earlier work this paper cites.
J. Dai, Y. Li, K. He, and J. Sun, “R-fcn: Object detection via region-based fully convolutional networks,” NeurIPS , 2016
2016
Earlier work this paper cites.
Y. Chebotar and A. Waters, “Distilling knowledge from ensembles of neural networks for speech recognition.” in Interspeech , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Gupta, J. Hoffman, and J. Malik, “Cross modal distillation for supervision transfer,” in CVPR , 2016
2016
Earlier work this paper cites.
N. Papernot, P. McDaniel, X. Wu, S. Jha, and A. Swami, “Distillation as a defense to adversarial perturbations against deep neural networks,” in SP , 2016
2016
Earlier work this paper cites.
M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang, “Deep learning with differential privacy,” in ACM SIGSAC , 2016
2016
Earlier work this paper cites.
P. Luo, Z. Zhu, Z. Liu, X. Wang, and X. Tang, “Face model compression by distilling knowledge from neurons,” in AAAI , 2016
2016
Earlier work this paper cites.
Y. Kim, Rush, and A. M, “Sequence-level knowledge distillation,” in EMNLP , 2016
2016
Earlier work this paper cites.
J. Yim, D. Joo, J. Bae, and J. Kim, “A gift from knowledge distillation: Fast optimization, network minimization and transfer learning,” in CVPR , 2017
2017
Earlier work this paper cites.
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska et al. , “Overcoming catastrophic forgetting in neural networks,” Proceedings of the national academy of sciences , 2017
2017
Earlier work this paper cites.
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in ICCV , 2017
2017
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, X. Zhang, and J. Sun, “Object detection networks on convolutional feature maps,” TPAMI , 2017
2017
Earlier work this paper cites.
G. Chen, W. Choi, X. Yu, T. Han, and M. Chandraker, “Learning efficient object detection models with knowledge distillation,” NeurIPS , 2017
2017
Earlier work this paper cites.
K. Shmelkov, C. Schmid, and K. Alahari, “Incremental learning of object detectors without catastrophic forgetting,” in ICCV , 2017
2017
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis, “Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer,” in ICLR , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. You, C. Xu, C. Xu, and D. Tao, “Learning from multiple teacher networks,” in KDD , 2017
2017
Earlier work this paper cites.
A. Tarvainen and H. Valpola, “Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results,” NeurIPS , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Z. Li and D. Hoiem, “Learning without forgetting,” in TPAMI , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
T. Asami, R. Masumura, Y. Yamaguchi, H. Masataki, and Y. Aono, “Domain adaptation of DNN acoustic models using knowledge distillation,” in ICASSP . IEEE, 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
X. Wang, R. Zhang, Y. Sun, and J. Qi, “Kdgan: Knowledge distillation with generative adversarial networks.” in NeurIPS , 2018
2018
Earlier work this paper cites.
J. Tang and K. Wang, “Ranking distillation: Learning compact ranking models with high performance for recommender system,” in KDD , 2018
2018
Earlier work this paper cites.
Y. Zhang, T. Xiang, T. M. Hospedales, and H. Lu, “Deep mutual learning,” in CVPR , 2018
2018
Earlier work this paper cites.
T. Furlanello, Z. Lipton, M. Tschannen, L. Itti, and A. Anandkumar, “Born again neural networks,” in ICML . PMLR, 2018
2018
Earlier work this paper cites.
Y.-H. Tsai, W.-C. Hung, S. Schulter, K. Sohn, M.-H. Yang, and M. Chandraker, “Learning to adapt structured output space for semantic segmentation,” in CVPR , 2018
2018
Earlier work this paper cites.
J. Hoffman, E. Tzeng, T. Park, J.-Y. Zhu, P. Isola, K. Saenko, A. Efros, and T. Darrell, “Cycada: Cycle-consistent adversarial domain adaptation,” in ICML . PMLR, 2018
2018
Earlier work this paper cites.
X. Zhu, S. Gong et al. , “Knowledge distillation by on-the-fly native ensemble,” NeurIPS , 2018
2018
Earlier work this paper cites.
M. Zhao, T. Li, M. Abu Alsheikh, Y. Tian, H. Zhao, A. Torralba, and D. Katabi, “Through-wall human pose estimation using radio signals,” in CVPR , 2018
2018
Earlier work this paper cites.
J. Li, R. Zhao, Z. Chen, C. Liu, X. Xiao, G. Ye, and Y. Gong, “Developing far-field speaker system via teacher-student learning,” in ICASSP . IEEE, 2018
2018
Earlier work this paper cites.
S. Ghorbani, A. E. Bulut, and J. H. Hansen, “Advancing multi-accented lstm-ctc speech recognition using a domain specific student-teacher learning paradigm,” in IEEE SLT , 2018
2018
Earlier work this paper cites.
J. Kim, S. Park, and N. Kwak, “Paraphrasing complex network: Network compression via factor transfer,” in NeurIPS , 2018
2018
Earlier work this paper cites.
Y. Chen, N. Wang, and Z. Zhang, “Darkrank: Accelerating deep metric learning via cross sample similarities transfer,” in AAAI , 2018
2018
Earlier work this paper cites.
N. Passalis and A. Tefas, “Learning deep representations with probabilistic knowledge transfer,” in ECCV , 2018
2018
Earlier work this paper cites.
Y. Wang, C. Xu, C. Xu, and D. Tao, “Adversarial Learning of Portable Student Networks,” in AAAI , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
N. C. Garcia, P. Morerio, and V. Murino, “Modality distillation with multiple stream networks for action recognition.” Springer, 2018
2018
Earlier work this paper cites.
H. Bagherinezhad, M. Horton, M. Rastegari, and A. Farhadi, “Label refinery: Improving imagenet classification through label progression,” 2018
2018
Earlier work this paper cites.
S. Albanie, A. Nagrani, A. Vedaldi, and A. Zisserman, “Emotion recognition in speech using cross-modal transfer in the wild,” in ACM MM , 2018
2018
Earlier work this paper cites.
S. Roheda, B. Riggan, H. Krim, and L. Dai, “Cross-modality distillation: A case for conditional generative adversarial networks,” in ICASSP , 2018
2018
Earlier work this paper cites.
S. Ghorbani, A. Bulut, and J. Hansen, “Advancing multi-accented lstm-ctc speech recognition using a domain specific student-teacher learning paradigm,” in SLTW , 2018
2018
Earlier work this paper cites.
C. Liu, B. Zoph, M. Neumann, J. Shlens, W. Hua, L.-J. Li, L. Fei-Fei, A. Yuille, J. Huang, and K. Murphy, “Progressive neural architecture search,” in ECCV , 2018
2018
Earlier work this paper cites.
T. Matiisen, A. Oliver, T. Cohen, and J. Schulman, “Teacher–student curriculum learning,” TNNLS , 2019
2019
Earlier work this paper cites.
Y. Li, L. Yuan, and N. Vasconcelos, “Bidirectional learning for domain adaptation of semantic segmentation,” in CVPR , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
W. Park, D. Kim, Y. Lu, and M. Cho, “Relational knowledge distillation,” in CVPR , 2019
2019
Cited alongside, same era.
F. Zhang, X. Zhu, and M. Ye, “Fast human pose estimation,” in CVPR , 2019
2019
Cited alongside, same era.
X. Nie, Y. Li, L. Luo, N. Zhang, and J. Feng, “Dynamic kernel distillation for efficient pose estimation in videos,” in ICCV , 2019
2019
Cited alongside, same era.
Z. Meng, J. Li, Y. Zhao, and Y. Gong, “Conditional teacher-student learning,” in ICASSP . IEEE, 2019
2019
Cited alongside, same era.
B. Shi, M. Sun, C.-C. Kao, V. Rozgic, S. Matsoukas, and C. Wang, “Semi-supervised acoustic event detection based on tri-training,” in ICASSP , 2019
2019
Cited alongside, same era.
H. Bai, J. Wu, I. King, and M. Lyu, “Few shot network compression via cross distillation,” in AAAI , 2020
2020
Later among the works it cites.
X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, and L. Li, “Tinybert: Distilling bert for natural language understanding,” in EMNLP , 2020
2020
Later among the works it cites.
P. Zhou, L. Mai, J. Zhang, N. Xu, Z. Wu, and L. Davis, “M2KD: Multi-model and multi-level knowledge distillation for incremental learning,” 2020
2020
Later among the works it cites.
J. Zhu, J. Liu, W. Li, J. Lai, X. He, L. Chen, and Z. Zheng, “Ensembled ctr prediction via knowledge distillation,” in CIKM , 2020
2020
Later among the works it cites.
S. Kang, J. Hwang, W. Kweon, and H. Yu, “DE-RRD: A knowledge distillation framework for recommender system,” in CIKM . ACM, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
B. Peng, X. Jin, J. Liu, D. Li, Y. Wu, Y. Liu, S. Zhou, and Z. Zhang, “Correlation congruence for knowledge distillation,” in ICCV , 2019
2019
Cited alongside, same era.
Y. Liu, J. Cao, B. Li, C. Yuan, W. Hu, Y. Li, and Y. Duan, “Knowledge distillation via instance relationship graph,” in CVPR , 2019
2019
Cited alongside, same era.
S. Ahn, S. X. Hu, A. Damianou, N. D. Lawrence, and Z. Dai, “Variational information distillation for knowledge transfer,” in CVPR , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
L. Zhang, J. Song, A. Gao, J. Chen, C. Bao, and K. Ma, “Be your own teacher: Improve the performance of convolutional neural networks via self distillation,” in ICCV , 2019
2019
Cited alongside, same era.
A. Wu, W.-S. Zheng, X. Guo, and J.-H. Lai, “Distilled person re-identification: Towards a more scalable system,” in CVPR , 2019
2019
Cited alongside, same era.
D. Liu, P. Cheng, Z. Dong, X. He, W. Pan, and Z. Ming, “A general knowledge distillation framework for counterfactual recommendation via uniform data,” in SIGIR , 2020
2020
Later among the works it cites.
M. Takamoto, Y. Morishita, and H. Imaoka, “An efficient method of training small models for regression problems with knowledge distillation,” in 2020 IEEE MIPR , 2020, pp. 67–72
2020
Later among the works it cites.
M. Kang, J. Mun, and B. Han, “Towards oracle knowledge distillation with neural architecture search,” in AAAI , 2020
2020
Later among the works it cites.
Y. Wang, Y. Yang, Y. Chen, J. Bai, C. Zhang, G. Su, X. Kou, Y. Tong, M. Yang, and L. Zhou, “Textnas: A neural architecture search space tailored for text representation,” in AAAI , 2020
2020
Later among the works it cites.
A. Mehrotra, A. G. C. Ramos, S. Bhattacharya, Ł. Dudziak, R. Vipperla, T. Chau, M. S. Abdelfattah, S. Ishtiaq, and N. D. Lane, “Nas-bench-asr: Reproducible neural architecture search for speech recognition,” in ICLR , 2020
2020
Later among the works it cites.
X. Cheng, Z. Rao, Y. Chen, and Q. Zhang, “Explaining knowledge distillation by quantifying the knowledge,” in CVPR , 2020
2020
Later among the works it cites.
S. Minaee, Y. Y. Boykov, F. Porikli, A. J. Plaza, N. Kehtarnavaz, and D. Terzopoulos, “Image segmentation using deep learning: A survey,” TPAMI , 2021
2021
Later among the works it cites.
Z. Wang, Y. Li, Y. Guo, L. Fang, and S. Wang, “Data-uncertainty guided multi-phase learning for semi-supervised object detection,” in CVPR , 2021
2021
Later among the works it cites.
G. Ghiasi, B. Zoph, E. D. Cubuk, Q. V. Le, and T.-Y. Lin, “Multi-task self-training for learning general representations,” in ICCV , 2021
2021
Later among the works it cites.
J. Gou, B. Yu, S. J. Maybank, and D. Tao, “Knowledge Distillation: A Survey,” in IJCV , 2021
2021
Later among the works it cites.
L. Wang and K.-J. Yoon, “Knowledge distillation and student-teacher learning for visual intelligence: A review and new outlooks,” TPAMI , 2021
2021
Later among the works it cites.
A. Alkhulaifi, F. Alsahli, and I. Ahmad, “Knowledge distillation in deep learning and its applications,” PeerJ Computer Science , 2021
2021
Later among the works it cites.
S. Stanton, P. Izmailov, P. Kirichenko, A. A. Alemi, and A. G. Wilson, “Does knowledge distillation really work?” NeurIPS , 2021
2021
Later among the works it cites.
Z. Li, J. Ye, M. Song, Y. Huang, and Z. Pan, “Online knowledge distillation for efficient pose estimation,” in ICCV , 2021
2021
Later among the works it cites.
J. Wang, S. Jin, W. Liu, W. Liu, C. Qian, and P. Luo, “When human pose estimation meets robustness: Adversarial algorithms and benchmarks,” in CVPR , 2021
2021
Later among the works it cites.
A. Chawla, H. Yin, P. Molchanov, and J. Alvarez, “Data-free knowledge distillation for object detection,” in WACV , 2021
2021
Later among the works it cites.
X. Tian, Z. Zhang, S. Lin, Y. Qu, Y. Xie, and L. Ma, “Farewell to mutual information: Variational distillation for cross-modal person re-identification,” in CVPR , 2021
2021
Later among the works it cites.
J. Zhu, S. Tang, D. Chen, S. Yu, Y. Liu, M. Rong, A. Yang, and X. Wang, “Complementary relation contrastive distillation,” in CVPR , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
L. T. Nguyen, K. Lee, and B. Shim, “Stochasticity and skip connection improve knowledge transfer,” in EUSIPCO , 2021
2021
Later among the works it cites.
F. Yuan, L. Shou, J. Pei, W. Lin, M. Gong, Y. Fu, and D. Jiang, “Reinforced multi-teacher selection for knowledge distillation,” in AAAI , 2021
2021
Later among the works it cites.
S. Zhou, Y. Wang, D. Chen, J. Chen, X. Wang, C. Wang, and J. Bu, “Distilling holistic knowledge with graph neural networks,” in ICCV , 2021
2021
Later among the works it cites.
C. Yang, J. Liu, and C. Shi, “Extract the knowledge of graph neural networks and go beyond it: An effective knowledge distillation framework,” in The WebConf , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
F. Sattler, T. Korjakow, R. Rischke, and W. Samek, “Fedaux: Leveraging unlabeled auxiliary data in federated learning,” TNNLS , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
F. Sattler, A. Marban, R. Rischke, and W. Samek, “Cfd: Communication-efficient federated distillation via soft-label quantization and delta coding,” IEEE Trans. Netw. Sci. Eng. , 2021
2021
Later among the works it cites.
L. Hu, H. Yan, L. Li, Z. Pan, X. Liu, and Z. Zhang, “Mhat: An efficient model-heterogenous aggregation training scheme for federated learning,” Information Sciences , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
S. Itahara, T. Nishio, Y. Koda, M. Morikura, and K. Yamamoto, “Distillation-based semi-supervised federated learning for communication-efficient collaborative training with non-iid private data,” IEEE Trans. Mobile Comput. , 2021
2021
Later among the works it cites.
H. Nguyen et al. , “Multimodal knowledge distillation for video and language processing,” TPAMI , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Kim, M. Hyun, I. Chung, and N. Kwak, “Feature fusion for online mutual knowledge distillation,” in ICPR . IEEE, 2021
2021
Later among the works it cites.
T. Su, Q. Liang, J. Zhang, Z. Yu, G. Wang, and X. Liu, “Attention-based feature interaction for efficient online knowledge distillation,” in ICDM . IEEE, 2021
2021
Later among the works it cites.
G. Wu and S. Gong, “Peer collaborative learning for online knowledge distillation,” in AAAI , 2021
2021
Later among the works it cites.
T. Isobe, X. Jia, S. Chen, J. He, Y. Shi, J. Liu, H. Lu, and S. Wang, “Multi-target domain adaptation with collaborative consistency learning,” in CVPR , 2021
2021
Later among the works it cites.
Z. Xue, S. Ren, Z. Gao, and H. Zhao, “Multimodal knowledge expansion,” in ICCV , 2021
2021
Later among the works it cites.
S. Chelaramani, M. Gupta, V. Agarwal, P. Gupta, and R. Habash, “Multi-task knowledge distillation for eye disease prediction,” in WACV , 2021
2021
Later among the works it cites.
M. Tzelepi, N. Passalis, and A. Tefas, “Online subclass knowledge distillation,” Expert Systems with Applications , 2021
2021
Later among the works it cites.
S. Kang, J. Hwang, W. Kweon, and H. Yu, “Topology distillation for recommender system,” in KDD , 2021
2021
Later among the works it cites.
M. Kang and S. Kang, “Data-free knowledge distillation in neural networks for regression,” Expert Syst. Appl. , 2021
2021
Later among the works it cites.
J. Yang, B. Martínez, A. Bulat, and G. Tzimiropoulos, “Knowledge distillation via softmax regression representation learning,” in ICLR , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Kang and S. Kang, “Data-free knowledge distillation in neural networks for regression,” Expert Systems with Applications , 2021
2021
Later among the works it cites.
2022
Later among the works it cites.
L. Beyer, X. Zhai, A. Royer, L. Markeeva, R. Anil, and A. Kolesnikov, “Knowledge distillation: A good teacher is patient and consistent,” in CVPR , 2022
2022
Later among the works it cites.
B. Zhao, Q. Cui, R. Song, Y. Qiu, and J. Liang, “Decoupled knowledge distillation,” in CVPR , 2022
2022
Later among the works it cites.
Z. Zheng, R. Ye, P. Wang, D. Ren, W. Zuo, Q. Hou, and M.-M. Cheng, “Localization distillation for dense object detection,” in CVPR , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
C. Yang, H. Zhou, Z. An, X. Jiang, Y. Xu, and Q. Zhang, “Cross-image relational knowledge distillation for semantic segmentation,” in CVPR , 2022
2022
Later among the works it cites.
C. K. Joshi, F. Liu, X. Xun, J. Lin, and C. S. Foo, “On representation knowledge distillation for graph neural networks,” TNNLS , 2022
2022
Later among the works it cites.
K. Feng, C. Li, Y. Yuan, and G. Wang, “Freekd: Free-direction knowledge distillation for graph neural networks,” in KDD , 2022
2022
Later among the works it cites.
Y. He, Y. Chen, X. Yang, H. Yu, Y.-H. Huang, and Y. Gu, “Learning critically: Selective self-distillation in federated learning on non-iid data,” IEEE Trans. Big Data , 2022
2022
Later among the works it cites.
Y. He, Y. Chen, X. Yang, Y. Zhang, and B. Zeng, “Class-wise adaptive self distillation for heterogeneous federated learning,” in AAAI , 2022
2022
Later among the works it cites.
L. Zhang, L. Shen, L. Ding, D. Tao, and L.-Y. Duan, “Fine-tuning global model via data-free knowledge distillation for non-iid federated learning,” in CVPR , 2022
2022
Later among the works it cites.
Q. Xu, Z. Chen, M. Ragab, C. Wang, M. Wu, and X. Li, “Contrastive adversarial knowledge distillation for deep model compression in time-series regression tasks,” Neurocomputing , 2022
2022
Later among the works it cites.
Q. Xu, Z. Chen, M. Ragab, C. Wang, M. Wu, and X. Li, “Contrastive adversarial knowledge distillation for deep model compression in time-series regression tasks,” Neurocomputing , 2022
2022
Later among the works it cites.
A. Shrivastava, Y. Qi, and V. Ordonez, “Estimating and maximizing mutual information for knowledge distillation,” in CVPR Workshops , 2023
2023
Closest in time.
G. Chen, J. Chen, F. Feng, S. Zhou, and X. He, “Unbiased knowledge distillation for recommendation,” in WSDM . ACM, 2023
2023
Closest in time.