Fetching the paper…
Reading the bibliography…
Knowledge distillation (KD) is a successful approach for deep neural network acceleration, with which a compact network (student) is trained by mimicking the softmax output of a pre-trained high-capacity network (teacher).
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Model compression
Buciluǎ, C., Caruana, R., and Niculescu-Mizil, A · 2006
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Do deep nets really need to be deep?
Ba, J. and Caruana, R · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
Jaderberg, M., Vedaldi, A., and Zisserman, A · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y · 2014
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Deep neural networks for youtube recommendations
Covington, P., Adams, J., and Sargin, E · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Pruning filters for efficient convnets
Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H. P · 2016
Earlier work this paper cites.
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Decision-based adversarial attacks: Reliable attacks against black-box machine learning models
Brendel, W., Rauber, J., and Bethge, M · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H · 2017
Cited alongside, same era.
Data-free knowledge distillation for deep neural networks
Lopes, R. G., Fenu, S., and Starner, T · 2017
Cited alongside, same era.
Alexa vs. siri vs. cortana vs. google assistant: a comparison of speech-based natural user interfaces
López, G., Quesada, L., and Guerrero, L. A · 2017
Cited alongside, same era.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Zero-shot knowledge transfer via adversarial belief matching
Micaelli, P. and Storkey, A. J · 2019
Later among the works it cites.
Zero-shot knowledge distillation in deep networks
Nayak, G. K., Mopuri, K. R., Shaj, V., Babu, R. V., and Chakraborty, A · 2019
Later among the works it cites.
Knockoff nets: Stealing functionality of black-box models
Orekondy, T., Schiele, B., and Fritz, M · 2019
Later among the works it cites.
Towards understanding knowledge distillation
Phuong, M. and Lampert, C · 2019
Later among the works it cites.
Clarifai: Computer vision and ai enterprise platform
Clarifai, I · 2020
Later among the works it cites.
Uncertainty-aware multi-shot knowledge distillation for image-based object re-identification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yim, J., Joo, D., Bae, J., and Kim, J · 2017
Cited alongside, same era.
Query-efficient hard-label black-box attack: An optimization-based approach
Cheng, M., Le, T., Chen, P.-Y., Yi, J., Zhang, H., and Hsieh, C.-J · 2018
Cited alongside, same era.
Born again neural networks
Furlanello, T., Lipton, Z., Tschannen, M., Itti, L., and Anandkumar, A · 2018
Cited alongside, same era.
Paraphrasing complex network: Network compression via factor transfer
Kim, J., Park, S., and Kwak, N · 2018
Cited alongside, same era.
Few-shot learning of neural networks from scratch by pseudo example optimization
Kimura, A., Ghahramani, Z., Takeuchi, K., Iwata, T., and Ueda, N · 2018
Cited alongside, same era.
Model and training methods of autonomous navigation system for compact drones
Moskalenko, V., Moskalenko, A., Korobov, A., Boiko, O., Martynenko, S., and Borovenskyi, O · 2018
Cited alongside, same era.
Stochastic zeroth-order optimization in high dimensions
Wang, Y., Du, S., Balakrishnan, S., and Singh, A · 2018
Cited alongside, same era.
Data-free learning of student networks
Chen, H., Wang, Y., Xu, C., Yang, Z., Liu, C., Shi, B., Xu, C., Xu, C., and Tian, Q · 2019
Cited alongside, same era.
Jin, X., Lan, C., Zeng, W., and Chen, Z · 2020
Later among the works it cites.
Few sample knowledge distillation for efficient network compression
Li, T., Li, J., Liu, Z., and Zhang, C · 2020
Later among the works it cites.
Heterogeneous knowledge distillation using information flow modeling
Passalis, N., Tzelepi, M., and Tefas, A · 2020
Later among the works it cites.
Neural networks are more productive teachers than human raters: Active mixup for data-efficient knowledge distillation from a blackbox model
Wang, D., Li, Y., Wang, L., and Gong, B · 2020
Later among the works it cites.
Knowledge distillation meets self-supervision
Xu, G., Liu, Z., Li, X., and Loy, C. C · 2020
Later among the works it cites.
Dreaming to distill: Data-free knowledge transfer via deepinversion
Yin, H., Molchanov, P., Alvarez, J. M., Li, Z., Mallya, A., Hoiem, D., Jha, N. K., and Kautz, J · 2020
Later among the works it cites.
Revisiting knowledge distillation via label smoothing regularization
Yuan, L., Tay, F. E., Li, G., Wang, T., and Feng, J · 2020
Later among the works it cites.
Regularizing class-wise predictions via self-knowledge distillation
Yun, S., Park, J., Lee, K., and Shin, J · 2020
Later among the works it cites.
Data-free knowledge distillation with soft targeted transfer set synthesis
Wang, Z · 2021
Closest in time.
Convolutional neural network pruning with structural redundancy reduction
Wang, Z., Li, C., and Wang, X · 2021
Closest in time.