Fetching the paper…
Reading the bibliography…
Most existing distillation methods ignore the flexible role of the temperature in the loss function and fix it as a hyper-parameter that can be decided by an inefficient grid search.
Competence-based curriculum learning for neural machine translation
Platanios, E. A.; Stretcu, O.; Neubig, G.; Poczos, B.; and Mitchell, T. M. 2019 · 1903
Earlier work this paper cites.
Tay, Y.; Wang, S.; Tuan, L. A.; Fu, J.; Phan, M. C.; Yuan, X.; Rao, J.; Hui, S. C.; and Zhang, A. 2019 · 1905
Earlier work this paper cites.
Caubrière, A.; Tomashenko, N.; Laurent, A.; Morin, E.; Camelin, N.; and Estève, Y. 2019 · 1906
Earlier work this paper cites.
Contrastive representation distillation
Tian, Y.; Krishnan, D.; and Isola, P. 2019 · 1910
Earlier work this paper cites.
An empirical analysis of the impact of data augmentation on knowledge distillation
Das, D.; Massa, H.; Kulkarni, A.; and Rekatsinas, T. 2020 · 2006
Earlier work this paper cites.
Curriculum learning
Bengio, Y.; Louradour, J.; Collobert, R.; and Weston, J. 2009 · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020 · 2010
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A.; Ballas, N.; Kahou, S. E.; Chassang, A.; Gatta, C.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K.; and Zisserman, A. 2014 · 2014
Earlier work this paper cites.
Unsupervised domain adaptation by backpropagation
Ganin, Y.; and Lempitsky, V. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Zagoruyko, S.; and Komodakis, N. 2016 · 2016
Earlier work this paper cites.
Reverse curriculum generation for reinforcement learning
Florensa, C.; Held, D.; Wulfmeier, M.; Zhang, M.; and Abbeel, P. 2017 · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; and Adam, H. 2017 · 2017
Cited alongside, same era.
Progressive growing of gans for improved quality, stability, and variation
Karras, T.; Aila, T.; Laine, S.; and Lehtinen, J. 2017 · 2017
Cited alongside, same era.
Feature pyramid networks for object detection
Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; and Belongie, S. 2017 · 2017
Cited alongside, same era.
Curriculum dropout
Morerio, P.; Cavazza, J.; Volpi, R.; Vidal, R.; and Murino, V. 2017 · 2017
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019 · 2019
Later among the works it cites.
Correlation congruence for knowledge distillation
Peng, B.; Jin, X.; Liu, J.; Li, D.; Wu, Y.; Liu, Y.; Zhou, S.; and Zhang, Z. 2019 · 2019
Later among the works it cites.
Similarity-preserving knowledge distillation
Tung, F.; and Mori, G. 2019 · 2019
Later among the works it cites.
Student becoming the master: Knowledge amalgamation for joint scene parsing, depth estimation, and more
Ye, J.; Ji, Y.; Wang, X.; Ou, K.; Tao, D.; and Song, M. 2019 · 2019
Later among the works it cites.
Knowledge distillation via instance-level sequence learning
Zhao, H.; Sun, X.; Dong, J.; Dong, Z.; and Li, Q. 2021 · 2019
Later among the works it cites.
Online Knowledge Distillation with Diverse Peers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Yim, J.; Joo, D.; Bae, J.; and Kim, J. 2017 · 2017
Cited alongside, same era.
Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels
Jiang, L.; Zhou, Z.; Leung, T.; Li, L.-J.; and Fei-Fei, L. 2018 · 2018
Cited alongside, same era.
Screenernet: Learning self-paced curriculum for deep neural networks
Kim, T.-H.; and Choi, J. 2018 · 2018
Cited alongside, same era.
Shufflenet v2: Practical guidelines for efficient cnn architecture design
Ma, N.; Zhang, X.; Zheng, H.-T.; and Sun, J. 2018 · 2018
Cited alongside, same era.
Learning deep representations with probabilistic knowledge transfer
Passalis, N.; and Tefas, A. 2018 · 2018
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; and Chen, L.-C. 2018 · 2018
Cited alongside, same era.
Learning to teach with dynamic loss functions
Wu, L.; Tian, F.; Xia, Y.; Fan, Y.; Qin, T.; Jian-Huang, L.; and Liu, T.-Y. 2018 · 2018
Cited alongside, same era.
Chen, D.; Mei, J.-P.; Wang, C.; Feng, Y.; and Chen, C. 2020 · 2020
Later among the works it cites.
Curriculum deepsdf
Duan, Y.; Zhu, H.; Wang, H.; Yi, L.; Nevatia, R.; and Guibas, L. J. 2020 · 2020
Later among the works it cites.
Curriculum by smoothing
Sinha, S.; Garg, A.; and Larochelle, H. 2020 · 2020
Later among the works it cites.
Learning from multiple experts: Self-paced knowledge distillation for long-tailed classification
Xiang, L.; Ding, G.; and Han, J. 2020 · 2020
Later among the works it cites.
Distilling knowledge via knowledge review
Chen, P.; Liu, S.; Zhao, H.; and Jia, J. 2021 · 2021
Later among the works it cites.
Refine Myself by Teaching Myself: Feature Refinement via Self-Knowledge Distillation
Ji, M.; Shin, S.; Hwang, S.; Park, G.; and Moon, I.-C. 2021 · 2021
Later among the works it cites.
A survey on curriculum learning
Wang, X.; Chen, Y.; and Zhu, W. 2021 · 2021
Later among the works it cites.
Knowledge distillation via softmax regression representation learning
Yang, J.; Martinez, B.; Bulat, A.; Tzimiropoulos, G.; et al. 2021 · 2021
Later among the works it cites.
Revisiting Label Smoothing and Knowledge Distillation Compatibility: What was Missing?
Chandrasegaran, K.; Tran, N.-T.; Zhao, Y.; and Cheung, N.-M. 2022 · 2022
Closest in time.
Knowledge distillation for object detection via rank mimicking and prediction-guided feature imitation
Li, G.; Li, X.; Wang, Y.; Zhang, S.; Wu, Y.; and Liang, D. 2022 · 2022
Closest in time.
Liu, J.; Liu, B.; Li, H.; and Liu, Y. 2022 · 2022
Closest in time.
Decoupled Knowledge Distillation
Zhao, B.; Cui, Q.; Song, R.; Qiu, Y.; and Liang, J. 2022 · 2022
Closest in time.