Fetching the paper…
Reading the bibliography…
Knowledge Distillation (KD) consists of transferring “knowledge†from one machine learning model (the teacher) to another (the student).
Neural network ensembles
Hansen, L. K. and Salamon, P · 1990
Earlier work this paper cites.
Society of mind: A response to four reviews
Minsky, M · 1991
Earlier work this paper cites.
Building a large annotated corpus of English: The penn treebank
Marcus, M. P., Marcinkiewicz, M. A., and Santorini, B · 1993
Earlier work this paper cites.
Born again trees
Breiman, L. and Shang, N · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Statistical modeling: The two cultures (with comments and a rejoinder by the author)
Breiman, L · 2001
Earlier work this paper cites.
Classification and regression by randomforest
Liaw, A., Wiener, M · 2002
Earlier work this paper cites.
Model compression
Bucilua, C., Caruana, R., and Niculescu-Mizil, A · 2006
Earlier work this paper cites.
The combination and comparison of neural networks with decision trees for wine classification
Chandra, R., Chaudhary, K., and Kumar, A · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Černockỳ, J., and Khudanpur, S · 2010
Earlier work this paper cites.
On the theory of learning with privileged information
Pechyony, D. and Vapnik, V · 2010
Earlier work this paper cites.
Do deep nets really need to be deep?
Ba, J. and Caruana, R · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Zaremba, W., Sutskever, I., and Vinyals, O · 2014
Earlier work this paper cites.
A neural algorithm of artistic style
Gatys, L. A., Ecker, A. S., and Bethge, M · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C. and Fei-Fei, L · 2015
Cited alongside, same era.
Deep learning, dark knowledge, and dark matter
Sadowski, P., Collado, J., Whiteson, D., and Baldi, P · 2015
Cited alongside, same era.
Xgboost: A scalable tree boosting system
Chen, T. and Guestrin, C · 2016
Cited alongside, same era.
Active long term memory networks
Furlanello, T., Zhao, J., Saxe, A. M., Itti, L., and Tjan, B. S · 2016
Cited alongside, same era.
Sequence-level knowledge distillation
Improved regularization of convolutional neural networks with cutout
DeVries, T. and Taylor, G. W · 2017
Later among the works it cites.
Coupled Ensembles of Neural Networks
Dutt, A., Pellerin, D., and Quenot, G · 2017
Later among the works it cites.
Distilling a neural network into a soft decision tree
Frosst, N. and Hinton, G · 2017
Later among the works it cites.
Gastaldi, X. · 2017
Later among the works it cites.
Deep pyramidal residual networks
Han, D., Kim, J., and Kim, J · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kim, Y. and Rush, A. M · 2016
Cited alongside, same era.
Character-aware neural language models
Kim, Y., Jernite, Y., Sontag, D., and Rush, A. M · 2016
Cited alongside, same era.
Learning without forgetting
Li, Z. and Hoiem, D · 2016
Cited alongside, same era.
The mythos of model interpretability
Lipton, Z. C · 2016
Cited alongside, same era.
Unifying distillation and privileged information
Lopez-Paz, D., Bottou, L., Schölkopf, B., and Vapnik, V · 2016
Cited alongside, same era.
Distillation as a defense to adversarial perturbations against deep neural networks
Papernot, N., McDaniel, P., Wu, X., Jha, S., and Swami, A · 2016
Cited alongside, same era.
Using the output embedding to improve language models
Press, O. and Wolf, L · 2016
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Liu, Z., Weinberger, K. Q., and van der Maaten, L · 2017
Later among the works it cites.
Snapshot ensembles: Train 1, get M for free
Huang, G., Li, Y., Pleiss, G., Liu, Z., Hopcroft, J. E., and Weinberger, K. Q · 2017
Later among the works it cites.
SGDR: Stochastic gradient descent with restarts
Loshchilov, I. and Hutter, F · 2017
Later among the works it cites.
Regularizing and optimizing LSTM language models
Merity, S., Keskar, N. S., and Socher, R · 2017
Later among the works it cites.
Continual learning with deep generative replay
Shin, H., Lee, J. K., Kim, J., and Kim, J · 2017
Later among the works it cites.
Do deep convolutional nets really need to be deep and convolutional?
Urban, G., Geras, K. J., Kahou, S. E., Aslan, O., Wang, S., Caruana, R., Mohamed, A., Philipose, M., and Richardson, M · 2017
Later among the works it cites.
In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pp. 5987–5995, 2017
Xie, S., Girshick, R., Dollár, P., Tu, Z., and He, K · 2017
Later among the works it cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Yim, J., Joo, D., Bae, J., and Kim, J · 2017
Later among the works it cites.
Transparent model distillation
Tan, S., Caruana, R., Hooker, G., and Gordo, A · 2018
Closest in time.
Yamada, Y., Iwamura, M., and Kise, K · 2018
Closest in time.