Fetching the paper…
Reading the bibliography…
Knowledge distillation deals with the problem of training a smaller model (Student) from a high capacity source model (Teacher) so as to retain most of its performance.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
A primer on statistical distributions
Balakrishnan, N. and Nevzorov, V. B · 2004
Earlier work this paper cites.
Model compression
Buciluǎ, C., Caruana, R., and Niculescu-Mizil, A · 2006
Earlier work this paper cites.
Sparse gaussian processes using pseudo-inputs
Snelson, E. and Ghahramani, Z · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Google deep dream
Mordvintsev, A., Tyka, M., and Olah, C · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Cited alongside, same era.
Striving for simplicity: The all convolutional net
Data-free knowledge distillation for deep neural networks
Lopes, R. G., Fenu, S., and Starner, T · 2017
Later among the works it cites.
Feature visualization
Olah, C., Mordvintsev, A., and Schubert, L · 2017
Later among the works it cites.
Fashion-MNIST: a novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Later among the works it cites.
A kernel theory of modern data augmentation
Dao, T., Gu, A., Ratner, A. J., Smith, V., De Sa, C., and Ré, C · 2018
Later among the works it cites.
Furlanello, T., Lipton, Z. C., Tschannen, M., Itti, L., and Anandkumar, A · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Springenberg, J., Dosovitskiy, A., Brox, T., and Riedmiller, M · 2015
Cited alongside, same era.
On the dirichlet distribution
Lin, J · 2016
Cited alongside, same era.
Few-shot learning of neural networks from scratch by pseudo example optimization
Kimura, A., Ghahramani, Z., Takeuchi, K., Iwata, T., and Ueda, N · 2018
Later among the works it cites.
Ask, acquire, and attack: Data-free uap generation using class impressions
Mopuri, K. R., Krishna, P., and Babu, R. V · 2018
Later among the works it cites.