Fetching the paper…
Reading the bibliography…
Knowledge distillation is a popular technique to transfer knowledge from large teacher models to a small student model.
Statistics and information theory, 1959
Solomon Kullback · 1959
Earlier work this paper cites.
Theory of classification: A survey of some recent advances
Stéphane Boucheron, Olivier Bousquet, and Gábor Lugosi · 2005
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al · 2015
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár · 2017
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Earlier work this paper cites.
Towards near-imperceptible steganographic text
Falcon Z. Dai and Zheng Jon Cai · 2019
Earlier work this paper cites.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E. Hinton · 2019
Cited alongside, same era.
Contrastive representation distillation
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2019
Cited alongside, same era.
Do we need zero training loss after achieving zero training error?
Takashi Ishida, Ikko Yamane, Tomoya Sakai, Gang Niu, and Masashi Sugiyama · 2020
Cited alongside, same era.
Knowledge distillation in wide neural networks: Risk bound, data efficiency and imperfect teacher
Guangda Ji and Zhanxing Zhu · 2020
Cited alongside, same era.
Adversarial nli: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 2020
Cited alongside, same era.
Rectifying the data bias in knowledge distillation
Boxiao Liu, Shenghan Zhang, Guanglu Song, Haihang You, and Yu Liu · 2021
Later among the works it cites.
Teacher’s pet: understanding and mitigating biases in distillation
Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, and Sanjiv Kumar · 2021
Later among the works it cites.
A statistical perspective on distillation
Aditya K Menon, Ankit Singh Rawat, Sashank Reddi, Seungyeon Kim, and Sanjiv Kumar · 2021
Later among the works it cites.
Does knowledge distillation really work?
Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A Alemi, and Andrew G Wilson · 2021
Later among the works it cites.
Rethinking soft labels for knowledge distillation: A bias-variance tradeoff perspective
Helong Zhou, Liangchen Song, Jiajie Chen, Ye Zhou, Guoli Wang, Junsong Yuan, and Qian Zhang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al · 2020
Cited alongside, same era.
Near-imperceptible neural linguistic steganography via self-adjusting arithmetic coding
Jiaming Shen, Heng Ji, and Jiawei Han · 2020
Cited alongside, same era.
Revisiting knowledge distillation via label smoothing regularization
Li Yuan, Francis EH Tay, Guilin Li, Tao Wang, and Jiashi Feng · 2020
Cited alongside, same era.
Distilling knowledge via knowledge review
Pengguang Chen, Shu Liu, Hengshuang Zhao, and Jiaya Jia · 2021
Cited alongside, same era.
Optimizing loss functions through multi-variate taylor polynomial parameterization
Santiago Gonzalez and Risto Miikkulainen · 2021
Cited alongside, same era.
Generalization bounds via distillation
Daniel Hsu, Ziwei Ji, Matus Telgarsky, and Lan Wang · 2021
Cited alongside, same era.
Polyloss: A polynomial expansion perspective of classification loss functions
Zhaoqi Leng, Mingxing Tan, Chenxi Liu, Ekin Dogus Cubuk, Jay Shi, Shuyang Cheng, and Dragomir Anguelov · 2022
Later among the works it cites.
Chatgpt, 2022
OpenAI · 2022
Later among the works it cites.
Better supervisory signals by observing learning paths
Yi Ren, Shangmin Guo, and Danica J Sutherland · 2022
Later among the works it cites.
Decoupled knowledge distillation
Borui Zhao, Quan Cui, Renjie Song, Yiyu Qiu, and Jiajun Liang · 2022
Later among the works it cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.