Fetching the paper…
Reading the bibliography…
This paper addresses the problem of model compression via knowledge distillation.
Model compression
Buciluǎ, C., Caruana, R., Niculescu-Mizil, A.: · 2006
Earlier work this paper cites.
Introduction to information retrieval (chapter 16)
Manning, C.D., Raghavan, P., Schütze, H.: · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., et al.: · 2009
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A., Ballas, N., Kahou, S.E., Chassang, A., Gatta, C., Bengio, Y.: · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., Lipson, H.: · 2014
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Han, S., Mao, H., Dally, W.J.: · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., Dean, J.: · 2015
Earlier work this paper cites.
Generative moment matching networks
Li, Y., Swersky, K., Zemel, R.: · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A.C., Fei-Fei, L.: · 2015
Earlier work this paper cites.
Fast convnets using group-wise brain damage
Lebedev, V., Lempitsky, V.: · 2016
Earlier work this paper cites.
XNOR-Net: ImageNet classification using binary convolutional neural networks
Rastegari, M., Ordonez, V., Redmon, J., Farhadi, A.: · 2016
Earlier work this paper cites.
Quantized convolutional neural networks for mobile devices
Wu, J., Leng, C., Wang, Y., Hu, Q., Cheng, J.: · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Zoph, B., Le, Q.V.: · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., Sun, J.: · 2016
Earlier work this paper cites.
Wide residual networks
Zagoruyko, S., Komodakis, N.: · 2016
Cited alongside, same era.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Zagoruyko, S., Komodakis, N.: · 2017
Cited alongside, same era.
Arbitrary style transfer in real-time with adaptive instance normalization
Huang, X., Belongie, S.: · 2017
Cited alongside, same era.
Revisiting batch normalization for practical domain adaptation
Li, Y., Wang, N., Shi, J., Liu, J., Hou, X.: · 2017
Cited alongside, same era.
Demystifying neural style transfer
Li, Y., Wang, N., Liu, J., Hou, X.: · 2017
Cited alongside, same era.
Like what you like: Knowledge distill via neuron selectivity transfer
Huang, Z., Wang, N.: · 2017
Cited alongside, same era.
Self-supervised knowledge distillation using singular value decomposition
Lee, S.H., Kim, D.H., Song, B.C.: · 2018
Later among the works it cites.
Paraphrasing complex network: Network compression via factor transfer
Kim, J., Park, S., Kwak, N.: · 2018
Later among the works it cites.
MobileNetV2: Inverted residuals and linear bottlenecks
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., Chen, L.C.: · 2018
Later among the works it cites.
On the efficacy of knowledge distillation
Cho, J.H., Hariharan, B.: · 2019
Later among the works it cites.
Knowledge transfer via distillation of activation boundaries formed by hidden neurons
Heo, B., Lee, M., Yun, S., Choi, J.Y.: · 2019
Later among the works it cites.
A comprehensive overhaul of feature distillation
Heo, B., Kim, J., Yun, S., Park, H., Kwak, N., Choi, J.Y.: · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Yim, J., Joo, D., Bae, J., Kim, J.: · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., Adam, H.: · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., Lerer, A.: · 2017
Cited alongside, same era.
Darts: Differentiable architecture search
Liu, H., Simonyan, K., Yang, Y.: · 2018
Cited alongside, same era.
On the power of over-parametrization in neural networks with quadratic activation
Du, S.S., Lee, J.D.: · 2018
Cited alongside, same era.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Soltanolkotabi, M., Javanmard, A., Lee, J.D.: · 2018
Cited alongside, same era.
Relational knowledge distillation
Park, W., Kim, D., Lu, Y., Cho, M.: · 2019
Later among the works it cites.
Correlation congruence for knowledge distillation
Peng, B., Jin, X., Liu, J., Zhou, S., Wu, Y., Liu, Y., Li, D., Zhang, Z.: · 2019
Later among the works it cites.
Knowledge distillation via instance relationship graph
Liu, Y., Cao, J., Li, B., Yuan, C., Hu, W., Li, Y., Duan, Y.: · 2019
Later among the works it cites.
Similarity-preserving knowledge distillation
Tung, F., Mori, G.: · 2019
Later among the works it cites.
Few-shot adversarial learning of realistic neural talking head models
Zakharov, E., Shysheya, A., Burkov, E., Lempitsky, V.: · 2019
Later among the works it cites.
AdaptIS: Adaptive instance selection network
Sofiiuk, K., Barinova, O., Konushin, A.: · 2019
Later among the works it cites.
XNOR-Net++: Improved binary neural networks
Bulat, A., Tzimiropoulos, G.: · 2019
Later among the works it cites.
Contrastive representation distillation
Tian, Y., Krishnan, D., Isola, P.: · 2020
Closest in time.