Fetching the paper…
Reading the bibliography…
Knowledge distillation (KD) has been extensively employed to transfer the knowledge from a large teacher model to the smaller students, where the parameters of the teacher are fixed (or partially) during training.
Supervised learning of probability distributions by neural networks
Eric B. Baum and Frank Wilczek. 1987 · 1987
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li. 2009 · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al. 2009 · 2009
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. 2016 · 2016
Earlier work this paper cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Junho Yim, Donggyu Joo, Ji-Hoon Bae, and Junmo Kim. 2017 · 2017
Earlier work this paper cites.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Sergey Zagoruyko and Nikos Komodakis. 2017 · 2017
Earlier work this paper cites.
SentEval: An evaluation toolkit for universal sentence representations
Alexis Conneau and Douwe Kiela. 2018 · 2018
Earlier work this paper cites.
Paraphrasing complex network: Network compression via factor transfer
Jangho Kim, Seonguk Park, and Nojun Kwak. 2018 · 2018
Earlier work this paper cites.
Learning deep representations with probabilistic knowledge transfer
Nikolaos Passalis and Anastasios Tefas. 2018 · 2018
Earlier work this paper cites.
Deep mutual learning
Ying Zhang, Tao Xiang, Timothy M. Hospedales, and Huchuan Lu. 2018 · 2018
Cited alongside, same era.
Variational information distillation for knowledge transfer
Sungsoo Ahn, Shell Xu Hu, Andreas C. Damianou, Neil D. Lawrence, and Zhenwen Dai. 2019 · 2019
Cited alongside, same era.
On the efficacy of knowledge distillation
Jang Hyun Cho and Bharath Hariharan. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Knowledge transfer via distillation of activation boundaries formed by hidden neurons
Byeongho Heo, Minsik Lee, Sangdoo Yun, and Jin Young Choi. 2019 · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Online knowledge distillation via collaborative learning
Qiushan Guo, Xinjiang Wang, Yichao Wu, Zhipeng Yu, Ding Liang, Xiaolin Hu, and Ping Luo. 2020 · 2020
Later among the works it cites.
Improved knowledge distillation via teacher assistant
Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh. 2020 · 2020
Later among the works it cites.
Contrastive representation distillation
Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2020 · 2020
Later among the works it cites.
BERT-of-theseus: Compressing BERT by progressive module replacing
Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou. 2020 · 2020
Later among the works it cites.
Distilling knowledge via knowledge review
Pengguang Chen, Shu Liu, Hengshuang Zhao, and Jiaya Jia. 2021 · 2021
Later among the works it cites.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen J. Maybank, and Dacheng Tao. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Knowledge distillation via route constrained optimization
Xiao Jin, Baoyun Peng, Yichao Wu, Yu Liu, Jiaheng Liu, Ding Liang, Junjie Yan, and Xiaolin Hu. 2019 · 2019
Cited alongside, same era.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey E. Hinton. 2019 · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E. Hinton. 2019 · 2019
Cited alongside, same era.
Relational knowledge distillation
Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho. 2019 · 2019
Cited alongside, same era.
Correlation congruence for knowledge distillation
Baoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou, Yichao Wu, Jiaheng Liu, Zhaoning Zhang, and Yu Liu. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Towards a unified view of parameter-efficient transfer learning
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. 2021 · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Later among the works it cites.
Learning student-friendly teacher networks for knowledge distillation
Dae Young Park, Moon-Hyun Cha, Daesin Kim, Bohyung Han, et al. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
Student can also be a good teacher: Extracting knowledge from vision-and-language model for cross-modal retrieval
Jun Rao, Tao Qian, Shuhan Qi, Yulin Wu, Qing Liao, and Xuan Wang. 2021 · 2021
Later among the works it cites.
Student customized knowledge distillation: Bridging the gap between student and teacher
Yichen Zhu and Yi Wang. 2021 · 2021
Later among the works it cites.
Reducing the teacher-student gap via adaptive temperatures
Jia Guo. 2022 · 2022
Closest in time.
Lora: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. 2022 · 2022
Closest in time.
Bert learns to teach: Knowledge distillation with meta learning
Wangchunshu Zhou, Canwen Xu, and Julian J. McAuley. 2022 · 2022
Closest in time.