Fetching the paper…
Reading the bibliography…
Several recent studies have elucidated why knowledge distillation (KD) improves model performance.
Captum: A unified and generic model interpretability library for pytorch
Kokhlikyan, N., Miglani, V., Martin, M., Wang, E., Alsallakh, B., Reynolds, J., Melnikov, A., Kliushkina, N., Araya, C., Yan, S., et al · 2009
Earlier work this paper cites.
Captum: A unified and generic model interpretability library for pytorch
Kokhlikyan, N., Miglani, V., Martin, M., Wang, E., Alsallakh, B., Reynolds, J., Melnikov, A., Kliushkina, N., Araya, C., Yan, S., et al · 2009
Earlier work this paper cites.
Intrinsic images in the wild
Bell, S., Bala, K., and Snavely, N · 2014
Earlier work this paper cites.
Detect what you can: Detecting and representing objects using holistic models and body parts
Chen, X., Mottaghi, R., Liu, X., Fidler, S., Urtasun, R., and Yuille, A · 2014
Earlier work this paper cites.
Describing textures in the wild
Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., and Vedaldi, A · 2014
Earlier work this paper cites.
The role of context for object detection and semantic segmentation in the wild
Mottaghi, R., Chen, X., Liu, X., Cho, N.-G., Lee, S.-W., Fidler, S., Urtasun, R., and Yuille, A · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y · 2014
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
” why should i trust you?” explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
Bau, D., Zhou, B., Khosla, A., Oliva, A., and Torralba, A · 2017
Earlier work this paper cites.
Real time image saliency for black box classifiers
Dabkowski, P. and Gal, Y · 2017
Earlier work this paper cites.
Interpretable explanations of black boxes by meaningful perturbation
Fong, R. C. and Vedaldi, A · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Earlier work this paper cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Yim, J., Joo, D., Bae, J., and Kim, J · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., and Torralba, A · 2017
Earlier work this paper cites.
Visualizing deep neural network decisions: Prediction difference analysis
Zintgraf, L. M., Cohen, T. S., Adel, T., and Welling, M · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A · 2017
Cited alongside, same era.
Towards robust interpretability with self-explaining neural networks
Alvarez-Melis, D. and Jaakkola, T. S · 2018
Cited alongside, same era.
This looks like that: deep learning for interpretable image recognition
Chen, C., Li, O., Tao, C., Barnett, A. J., Su, J., and Rudin, C · 2018
A learning theoretic perspective on local explainability
Li, J., Nagarajan, V., Plumb, G., and Talwalkar, A · 2020
Later among the works it cites.
Understanding and improving knowledge distillation
Tang, J., Shivanna, R., Zhao, Z., Lin, D., Singh, A., Chi, E. H., and Jain, S · 2020
Later among the works it cites.
Quantifying explainability of saliency methods in deep neural networks
Tjoa, E. and Guan, C · 2020
Later among the works it cites.
Knowledge distillation meets self-supervision
Xu, G., Liu, Z., Li, X., and Loy, C. C · 2020
Later among the works it cites.
Revisiting knowledge distillation via label smoothing regularization
Yuan, L., Tay, F. E., Li, G., Wang, T., and Feng, J · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Born again neural networks
Furlanello, T., Lipton, Z., Tschannen, M., Itti, L., and Anandkumar, A · 2018
Cited alongside, same era.
Paraphrasing complex network: Network compression via factor transfer
Kim, J., Park, S., and Kwak, N · 2018
Cited alongside, same era.
Robustness may be at odds with accuracy
Tsipras, D., Santurkar, S., Engstrom, L., Turner, A., and Madry, A · 2018
Cited alongside, same era.
Born again neural networks
Furlanello, T., Lipton, Z., Tschannen, M., Itti, L., and Anandkumar, A · 2018
Cited alongside, same era.
Paraphrasing complex network: Network compression via factor transfer
Kim, J., Park, S., and Kwak, N · 2018
Cited alongside, same era.
On the efficacy of knowledge distillation
Cho, J. H. and Hariharan, B · 2019
Cited alongside, same era.
Novelty detection and analysis in convolutional neural networks
Eshed, N · 2020
Later among the works it cites.
Quantifying explainability of saliency methods in deep neural networks
Tjoa, E. and Guan, C · 2020
Later among the works it cites.
Knowledge distillation meets self-supervision
Xu, G., Liu, Z., Li, X., and Loy, C. C · 2020
Later among the works it cites.
Bastings, J., Ebert, S., Zablotskaia, P., Sandholm, A., and Filippova, K · 2021
Later among the works it cites.
Reviewing the need for explainable artificial intelligence (xai)
Gerlings, J., Shollo, A., and Constantiou, I · 2021
Later among the works it cites.
The out-of-distribution problem in explainability and search methods for feature importance explanations
Hase, P., Xie, H., and Bansal, M · 2021
Later among the works it cites.
Enabling lightweight fine-tuning for pre-trained language model compression based on matrix product operators
Liu, P., Gao, Z.-F., Zhao, W. X., Xie, Z.-Y., Lu, Z.-Y., and Wen, J.-R · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Later among the works it cites.
Do input gradients highlight discriminative features?
Shah, H., Jain, P., and Netrapalli, P · 2021
Later among the works it cites.
Is label smoothing truly incompatible with knowledge distillation: An empirical study
Shen, Z., Liu, Z., Xu, D., Chen, Z., Cheng, K.-T., and Savvides, M · 2021
Later among the works it cites.
Symbolic knowledge distillation: from general language models to commonsense models
West, P., Bhagavatula, C., Hessel, J., Hwang, J. D., Jiang, L., Bras, R. L., Lu, X., Welleck, S., and Choi, Y · 2021
Later among the works it cites.
On the sensitivity and stability of model interpretations in nlp
Yin, F., Shi, Z., Hsieh, C.-J., and Chang, K.-W · 2021
Later among the works it cites.
Rethinking soft labels for knowledge distillation: A bias-variance tradeoff perspective
Zhou, H., Song, L., Chen, J., Zhou, Y., Wang, G., Yuan, J., and Zhang, Q · 2021
Later among the works it cites.
torchdistill: A modular, configuration-driven framework for knowledge distillation
Matsubara, Y · 2021
Later among the works it cites.
Revisiting label smoothing and knowledge distillation compatibility: What was missing?
Chandrasegaran, K., Tran, N.-T., Zhao, Y., and Cheung, N.-M · 2022
Later among the works it cites.
Explainable artificial intelligence (xai) in deep learning-based medical image analysis
van der Velden, B. H., Kuijf, H. J., Gilhuijs, K. G., and Viergever, M. A · 2022
Later among the works it cites.
Multimodal adaptive distillation for leveraging unimodal encoders for vision-language tasks
Wang, Z., Codella, N., Chen, Y.-C., Zhou, L., Dai, X., Xiao, B., Yang, J., You, H., Chang, K.-W., Chang, S.-f., et al · 2022
Later among the works it cites.