Fetching the paper…
Reading the bibliography…
Recently, large-scale pre-trained models have shown their advantages in many tasks.
B. Heo, J. Kim, S. Yun, H. Park, N. Kwak, and J. Y. Choi, “A comprehensive overhaul of feature distillation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 1921–1930
1930
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition , 2009, pp. 248–255
2009
Earlier work this paper cites.
L. Van Der Maaten and K. Weinberger, “Stochastic triplet embedding,” in 2012 IEEE International Workshop on Machine Learning for Signal Processing . IEEE, 2012, pp. 1–6
2012
Earlier work this paper cites.
J. Ba and R. Caruana, “Do deep nets really need to be deep,” in Advances in Neural Information Processing Systems 27 , vol. 27, 2014, pp. 2654–2662
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European conference on computer vision . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio, “Fitnets: Hints for thin deep nets,” in ICLR 2015 : International Conference on Learning Representations 2015 , 2015
2015
Earlier work this paper cites.
P. Molchanov, S. Tyree, T. Karras, T. Aila, and J. Kautz, “Pruning convolutional neural networks for resource efficient inference,” in ICLR (Poster) , 2016
2016
Earlier work this paper cites.
J. Wu, C. Leng, Y. Wang, Q. Hu, and J. Cheng, “Quantized convolutional neural networks for mobile devices,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 4820–4828
2016
Earlier work this paper cites.
Y. Chebotar and A. Waters, “Distilling knowledge from ensembles of neural networks for speech recognition.” in Interspeech , 2016, pp. 3439–3443
2016
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis, “Paying more attention to attention: improving the performance of convolutional neural networks via attention transfer,” in ICLR (Poster) , 2016
2016
Earlier work this paper cites.
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 3213–3223
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
M. H. Zhu and S. Gupta, “To prune, or not to prune: exploring the efficacy of pruning for model compression,” in ICLR (Workshop) , 2017
2017
Earlier work this paper cites.
Z. Huang and N. Wang, “Like what you like: Knowledge distill via neuron selectivity transfer.” arXiv: Computer Vision and Pattern Recognition , 2017
2017
Earlier work this paper cites.
J. Yim, D. Joo, J. Bae, and J. Kim, “A gift from knowledge distillation: Fast optimization, network minimization and transfer learning,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 7130–7138
2017
Earlier work this paper cites.
J. Snell, K. Swersky, and R. S. Zemel, “Prototypical networks for few-shot learning,” in Advances in Neural Information Processing Systems , vol. 30, 2017, pp. 4077–4087
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
H. Venkateswara, J. Eusebio, S. Chakraborty, and S. Panchanathan, “Deep hashing network for unsupervised domain adaptation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 5018–5027
2017
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. N. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , 2018, pp. 4171–4186
2018
Earlier work this paper cites.
N. Passalis and A. Tefas, “Learning deep representations with probabilistic knowledge transfer,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 283–299
2018
Cited alongside, same era.
G. Van Horn, O. Mac Aodha, Y. Song, Y. Cui, C. Sun, A. Shepard, H. Adam, P. Perona, and S. Belongie, “The inaturalist species classification and detection dataset,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 8769–8778
2018
Cited alongside, same era.
C. Sakaridis, D. Dai, and L. Van Gool, “Semantic foggy scene understanding with synthetic data,” International Journal of Computer Vision , vol. 126, no. 9, pp. 973–992, 2018
2018
Cited alongside, same era.
W. Choi, M. Chandraker, G. Chen, and X. Yu, “Learning efficient object detection models with knowledge distillation,” Sep. 20 2018, uS Patent App. 15/908,870
2018
Cited alongside, same era.
G. Kurata and G. Saon, “Knowledge distillation from offline to streaming rnn transducer for end-to-end speech recognition.” in Interspeech , 2020, pp. 2117–2121
2020
Later among the works it cites.
X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu, “Tinybert: Distilling bert for natural language understanding,” in Findings of the Association for Computational Linguistics: EMNLP 2020 , 2020, pp. 4163–4174
2020
Later among the works it cites.
H.-J. Ye, S. Lu, and D.-C. Zhan, “Distilling cross-task knowledge via relationship matching,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020, pp. 12 396–12 405
2020
Later among the works it cites.
I. Radosavovic, R. P. Kosaraju, R. Girshick, K. He, and P. Dollar, “Designing network design spaces,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020, pp. 10 428–10 436
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Ye, M. Rochan, Z. Liu, and Y. Wang, “Cross-modal self-attention network for referring image segmentation,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 10 502–10 511
2019
Cited alongside, same era.
H. Tan and M. Bansal, “Lxmert: Learning cross-modality encoder representations from transformers,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 5099–5110
2019
Cited alongside, same era.
R. R. Müller, S. Kornblith, and G. Hinton, “When does label smoothing help,” in Advances in Neural Information Processing Systems , vol. 32, 2019, pp. 4694–4703
2019
Cited alongside, same era.
W. Park, D. Kim, Y. Lu, and M. Cho, “Relational knowledge distillation,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 3967–3976
2019
Cited alongside, same era.
2019
Cited alongside, same era.
S. Sun, Y. Cheng, Z. Gan, and J. Liu, “Patient knowledge distillation for bert model compression,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 4322–4331
2019
Cited alongside, same era.
J. Liu, L. Song, and Y. Qin, “Prototype rectification for few-shot learning,” in European Conference on Computer Vision , 2019, pp. 741–756
2019
Cited alongside, same era.
2019
Cited alongside, same era.
L. Zhang and K. Ma, “Improve object detection with feature-based knowledge distillation: Towards accurate and efficient detectors,” in International Conference on Learning Representations , 2020
2020
Later among the works it cites.
H. Touvron, M. Cord, D. Matthijs, F. Massa, A. Sablayrolles, and H. Jegou, “Training data-efficient image transformers & distillation through attention,” in ICML 2021: 38th International Conference on Machine Learning , 2021, pp. 10 347–10 357
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 10 012–10 022
2021
Later among the works it cites.
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable detr: Deformable transformers for end-to-end object detection,” in ICLR 2021: The Ninth International Conference on Learning Representations , 2021
2021
Later among the works it cites.
Q. Zhao, J. Dong, H. Yu, and S. Chen, “Distilling ordinal relation and dark knowledge for facial age estimation,” IEEE Transactions on Neural Networks and Learning Systems , vol. 32, no. 7, pp. 3108–3121, 2021
2021
Later among the works it cites.
J. Gou, B. Yu, S. J. Maybank, and D. Tao, “Knowledge distillation: A survey,” International Journal of Computer Vision , vol. 129, no. 6, pp. 1789–1819, 2021
2021
Later among the works it cites.
J. W. Yoon, H. Lee, H. Y. Kim, W. I. Cho, and N. S. Kim, “Tutornet: Towards flexible knowledge distillation for end-to-end speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 1626–1638, 2021
2021
Later among the works it cites.
D. Chen, J.-P. Mei, Y. Zhang, C. Wang, Z. Wang, Y. Feng, and C. Chen, “Cross-layer distillation with semantic calibration,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 8, 2021, pp. 7028–7036
2021
Later among the works it cites.
H. Chefer, S. Gur, and L. Wolf, “Transformer interpretability beyond attention visualization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 782–791
2021
Later among the works it cites.
C. Shu, Y. Liu, J. Gao, Z. Yan, and C. Shen, “Channel-wise knowledge distillation for dense prediction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 5311–5320
2021
Later among the works it cites.
S. Alfasly, C. K. Chui, Q. Jiang, J. Lu, and C. Xu, “An effective video transformer with synchronized spatiotemporal and spatial self-attention for action recognition,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–14, 2022
2022
Closest in time.
S. Li, M. Lin, Y. Wang, Y. Wu, Y. Tian, L. Shao, and R. Ji, “Distilling a powerful student model via online knowledge distillation,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–10, 2022
2022
Closest in time.
M. Zhu, J. Li, N. Wang, and X. Gao, “Knowledge distillation for face photo–sketch synthesis,” IEEE Transactions on Neural Networks and Learning Systems , vol. 33, no. 2, pp. 893–906, 2022
2022
Closest in time.
C. Yang, Z. An, L. Cai, and Y. Xu, “Knowledge distillation using hierarchical self-supervision augmented distribution,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–15, 2022
2022
Closest in time.
L. Li, Y.-C. Chen, Y. Cheng, Z. Gan, L. Yu, and J. Liu, “Hero: Hierarchical encoder for video+language omni-representation pre-training,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 2046–2065
2065
Closest in time.