Fetching the paper…
Reading the bibliography…
Knowledge distillation is a popular technique for training a small student network to emulate a larger teacher model, such as an ensemble of networks.
Using a neural network to approximate an ensemble of classifiers
Zeng, X. and Martinez, T. R. (2000) · 2000
Earlier work this paper cites.
Scalable second order optimization for deep learning
Anil, R., Gupta, V., Koren, T., Regan, K., and Singer, Y. (2021) · 2002
Earlier work this paper cites.
Model compression
Bucilă, C., Caruana, R., and Niculescu-Mizil, A. (2006) · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. et al. (2009) · 2009
Earlier work this paper cites.
Torchvision the machine-vision package of torch
Marcel, S. and Rodriguez, Y. (2010) · 2010
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. (2011) · 2011
Earlier work this paper cites.
Do deep nets really need to be deep?
Ba, J. and Caruana, R. (2014) · 2014
Earlier work this paper cites.
Distilling knowledge from deep networks with applications to healthcare domain
Che, Z., Purushotham, S., Khemani, R., and Liu, Y. (2015) · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J. (2015) · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y. (2015) · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A. (2015) · 2015
Earlier work this paper cites.
Layer normalization
Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016) · 2016
Earlier work this paper cites.
Distilling knowledge from ensembles of neural networks for speech recognition
Chebotar, Y. and Waters, A. (2016) · 2016
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Han, S., Mao, H., and Dally, W. J. (2016) · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Deep neural networks with massive learned knowledge
Hu, Z., Yang, Z., Salakhutdinov, R., and Xing, E. (2016b) · 2016
Cited alongside, same era.
Sequence-level knowledge distillation
Kim, Y. and Rush, A. M. (2016) · 2016
Cited alongside, same era.
Distillation as a defense to adversarial perturbations against deep neural networks
Papernot, N., McDaniel, P., Wu, X., Jha, S., and Swami, A. (2016) · 2016
Cited alongside, same era.
Learning efficient object detection models with knowledge distillation
Chen, G., Choi, W., Yu, X., Han, T., and Chandraker, M. (2017) · 2017
Cited alongside, same era.
Emnist: Extending mnist to handwritten letters
Cohen, G., Afshar, S., Tapson, J., and Van Schaik, A. (2017) · 2017
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y. (2018) · 2018
Later among the works it cites.
Distill-and-compare: Auditing black-box models using transparent model distillation
Tan, S., Caruana, R., Hooker, G., and Lou, Y. (2018) · 2018
Later among the works it cites.
Between-class learning for image classification
Tokozume, Y., Ushiku, Y., and Harada, T. (2018) · 2018
Later among the works it cites.
Mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D. (2018) · 2018
Later among the works it cites.
On the efficacy of knowledge distillation
Cho, J. H. and Hariharan, B. (2019) · 2019
Later among the works it cites.
Knowledge transfer via distillation of activation boundaries formed by hidden neurons
Heo, B., Lee, M., Yun, S., and Choi, J. Y. (2019) · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017) · 2017
Cited alongside, same era.
Do deep convolutional nets really need to be deep and convolutional?
Urban, G., Geras, K. J., Kahou, S. E., Aslan, O., Wang, S., Caruana, R., Mohamed, A., Philipose, M., and Richardson, M. (2017) · 2017
Cited alongside, same era.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Yim, J., Joo, D., Bae, J., and Kim, J. (2017) · 2017
Cited alongside, same era.
Learning from multiple teacher networks
You, S., Xu, C., Xu, C., and Tao, D. (2017) · 2017
Cited alongside, same era.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Zagoruyko, S. and Komodakis, N. (2017) · 2017
Cited alongside, same era.
Born again neural networks
Furlanello, T., Lipton, Z., Tschannen, M., Itti, L., and Anandkumar, A. (2018) · 2018
Cited alongside, same era.
Later among the works it cites.
Similarity of neural network representations revisited
Kornblith, S., Norouzi, M., Lee, H., and Hinton, G. (2019) · 2019
Later among the works it cites.
Ensemble distribution distillation
Malinin, A., Mlodozeniec, B., and Gales, M. (2019) · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. (2019) · 2019
Later among the works it cites.
Meal: Multi-model ensemble via adversarial learning
Shen, Z., He, Z., and Xue, X. (2019) · 2019
Later among the works it cites.
Online knowledge distillation with diverse peers
Chen, D., Mei, J.-P., Wang, C., Feng, Y., and Chen, C. (2020) · 2020
Later among the works it cites.
Fast, accurate, and simple models for tabular data via augmented distillation
Fakoor, R., Mueller, J. W., Erickson, N., Chaudhari, P., and Smola, A. J. (2020) · 2020
Later among the works it cites.
Adversarially robust distillation
Goldblum, M., Fowl, L., Feizi, S., and Goldstein, T. (2020) · 2020
Later among the works it cites.
Self-distillation amplifies regularization in hilbert space
Mobahi, H., Farajtabar, M., and Bartlett, P. L. (2020) · 2020
Later among the works it cites.
Knowledge distillation: A good teacher is patient and consistent
Beyer, L., Zhai, X., Royer, A., Markeeva, L., Anil, R., and Kolesnikov, A. (2021) · 2021
Closest in time.