Fetching the paper…
Reading the bibliography…
Performing knowledge transfer from a large teacher network to a smaller student is a popular task in modern deep learning applications.
Variational information distillation for knowledge transfer
Ahn, S., Hu, S. X., Damianou, A. C., Lawrence, N. D., and Dai, Z. (2019) · 1904
Earlier work this paper cites.
Adversarial Examples Are Not Bugs, They Are Features
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A. (2019) · 1905
Earlier work this paper cites.
A unifying view of sparse approximate Gaussian process regression
Candela, J. Q. and Rasmussen, C. E. (2005) · 1959
Earlier work this paper cites.
Sparse Gaussian processes using pseudo-inputs
Snelson, E. and Ghahramani, Z. (2005) · 2005
Earlier work this paper cites.
Model compression
Buciluǎ, C., Caruana, R., and Niculescu-Mizil, A. (2006) · 2006
Earlier work this paper cites.
Zero-data learning of new tasks
Larochelle, H., Erhan, D., and Bengio, Y. (2008) · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. (2009) · 2009
Earlier work this paper cites.
Variational learning of inducing variables in sparse Gaussian processes
Titsias, M. K. (2009) · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. (2011) · 2011
Earlier work this paper cites.
Zero-shot learning through cross-modal transfer
Socher, R., Ganjoo, M., Manning, C. D., and Ng, A. Y. (2013) · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?
Ba, J. and Caruana, R. (2014) · 2014
Earlier work this paper cites.
Privacy in pharmacogenetics: An end-to-end case study of personalized warfarin dosing
Fredrikson, M., Lantz, E., Jha, S., Lin, S., Page, D., and Ristenpart, T. (2014) · 2014
Earlier work this paper cites.
Deepface: Closing the gap to human-level performance in face verification
Taigman, Y., Yang, M., Ranzato, M., and Wolf, L. (2014) · 2014
Earlier work this paper cites.
Data Safe Havens in health research and healthcare
Burton, P. R., Murtagh, M. J., Boyd, A., Williams, J. B., Dove, E. S., Wallace, S. E., Tassé, A.-M., Little, J., Chisholm, R. L., Gaye, A., Hveem, K., Brookes, A. J., Goodwin, P., Fistein, J., Bobrow, M., and Knoppers, B. M. (2015) · 2015
Cited alongside, same era.
Model inversion attacks that exploit confidence information and basic countermeasures
Fredrikson, M., Jha, S., and Ristenpart, T. (2015) · 2015
Cited alongside, same era.
Deep learning with limited numerical precision
Gupta, S., Agrawal, A., Gopalakrishnan, K., and Narayanan, P. (2015) · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2015) · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J. (2015) · 2015
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., Klingner, J., Shah, A., Johnson, M., Liu, X., Kaiser, L., Gouws, S., Kato, Y., Kudo, T., Kazawa, H., Stevens, K., Kurian, G., Patil, N., Wang, W., Young, C., Smith, J., Riesa, J., Rudnick, A., Vinyals, O., Corrado, G., Hughes, M., and Dean, J. (2016) · 2016
Later among the works it cites.
Dehghani, M., Mehrjou, A., Gouws, S., Kamps, J., and Schölkopf, B. (2017) · 2017
Later among the works it cites.
Data-free knowledge distillation for deep neural networks
Lopes, R. G., Fenu, S., and Starner, T. (2017) · 2017
Later among the works it cites.
Revisiting unreasonable effectiveness of data in deep learning era
Sun, C., Shrivastava, A., Singh, S., and Gupta, A. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Cited alongside, same era.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S. E., Chassang, A., Gatta, C., and Bengio, Y. (2015) · 2015
Cited alongside, same era.
Han, S., Mao, H., and Dally, W. J. (2016) · 2016
Cited alongside, same era.
Quantized neural networks: Training neural networks with low precision weights and activations
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y. (2016) · 2016
Cited alongside, same era.
Pruning filters for efficient convnets
Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H. P. (2016) · 2016
Cited alongside, same era.
Policy distillation
Rusu, A. A., Colmenarejo, S. G., Gülçehre, Ç., Desjardins, G., Kirkpatrick, J., Pascanu, R., Mnih, V., Kavukcuoglu, K., and Hadsell, R. (2016) · 2016
Cited alongside, same era.
Stealing machine learning models via prediction apis
Tramèr, F., Zhang, F., Juels, A., Reiter, M. K., and Ristenpart, T. (2016) · 2016
Cited alongside, same era.
Moonshine: Distilling with cheap convolutions
Crowley, E. J., Gray, G., and Storkey, A. (2018) · 2018
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K. (2018) · 2018
Later among the works it cites.
Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., and Brendel, W. (2018) · 2018
Later among the works it cites.
Few-shot learning of neural networks from scratch by pseudo example optimization
Kimura, A., Ghahramani, Z., Takeuchi, K., Iwata, T., and Ueda, N. (2018) · 2018
Later among the works it cites.
Knowledge distillation from few samples
Li, T., Li, J., Liu, Z., and Zhang, C. (2018) · 2018
Later among the works it cites.
Realistic evaluation of deep semi-supervised learning algorithms
Oliver, A., Odena, A., Raffel, C., Cubuk, E. D., and Goodfellow, I. J. (2018) · 2018
Later among the works it cites.
Wang, T., Zhu, J., Torralba, A., and Efros, A. A. (2018) · 2018
Later among the works it cites.
Zero-shot knowledge distillation in deep networks
Nayak, G. K., Mopuri, K. R., Shaj, V., Babu, R. V., and Chakraborty, A. (2019) · 2019
Closest in time.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. (2019) · 2019
Closest in time.