Variational information distillation for knowledge transfer
Ahn, S., Hu, S. X., Damianou, A. C., Lawrence, N. D., and Dai, Z. (2019) · 2019
Later among the works it cites.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F. (2019) · 2019
Later among the works it cites.
A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation
Gotmare, A., Keskar, N. S., Xiong, C., and Socher, R. (2019) · 2019
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J. (2019) · 2019
Later among the works it cites.
Zero-shot knowledge distillation in deep networks
Nayak, G. K., Mopuri, K. R., Shaj, V., Radhakrishnan, V. B., and Chakraborty, A. (2019) · 2019
Later among the works it cites.
Relational knowledge distillation
Park, W., Kim, D., Lu, Y., and Cho, M. (2019) · 2019
Later among the works it cites.
Towards understanding knowledge distillation
Phuong, M. and Lampert, C. (2019) · 2019
Later among the works it cites.
Training deep neural networks in generations: A more tolerant teacher educates better students
Yang, C., Xie, L., Qiao, S., and Yuille, A. L. (2019) · 2019
Later among the works it cites.
Transferring inductive biases through knowledge distillation
Abnar, S., Dehghani, M., and Zuidema, W. (2020) · 2020
Closest in time.
Collaborative inter-agent knowledge distillation for reinforcement learning
Hong, Z.-W., Nagarajan, P., and Maeda, G. (2020) · 2020
Closest in time.
Why distillation helps: a statistical perspective
Menon, A. K., Rawat, A. S., Reddi, S. J., Kim, S., and Kumar, S. (2020) · 2020
Closest in time.
Neural kernels without tangents
Shankar, V., Fang, A., Guo, W., Fridovich-Keil, S., Ragan-Kelley, J., Schmidt, L., and Recht, B. (2020) · 2020
Closest in time.
Hydra: Preserving ensemble diversity for model distillation
Tran, L., Veeling, B. S., Roth, K., Swiatkowski, J., Dillon, J. V., Snoek, J., Mandt, S., Salimans, T., Nowozin, S., and Jenatton, R. (2020) · 2020
Closest in time.
Generalized bayesian posterior expectation distillation for deep neural networks
Vadera, M. P. and Marlin, B. M. (2020) · 2020
Closest in time.