Fetching the paper…
Reading the bibliography…
Despite the popularity and efficacy of knowledge distillation, there is limited understanding of why it helps.
Probability in Banach Spaces: isoperimetry and processes , volume 23
Michel Ledoux and Michel Talagrand · 1991
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
On kernel-target alignment
Nello Cristianini, John Shawe-Taylor, Andre Elisseeff, and Jaz Kandola · 2001
Earlier work this paper cites.
Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond
Bernhard Scholkopf and Alexander J. Smola · 2001
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al · 2015
Earlier work this paper cites.
Rademacher complexity margin bounds for learning with a large number of classes
Vitaly Kuznetsov, Mehryar Mohri, and U Syed · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Samira Ebrahimi Kahou, Polytechnique Montréal, Y. Bengio, Université De Montréal, Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Sergey Zagoruyko and Nikos Komodakis · 2017
Earlier work this paper cites.
Large scale distributed neural network training through online distillation
Rohan Anil, Gabriel Pereyra, Alexandre Passos, Róbert Ormándi, George E. Dahl, and Geoffrey E. Hinton · 2018
Earlier work this paper cites.
To understand deep learning we need to understand kernel learning
Mikhail Belkin, Siyuan Ma, and Soumik Mandal · 2018
Earlier work this paper cites.
Born again neural networks
Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Earlier work this paper cites.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Earlier work this paper cites.
Learning deep representations with probabilistic knowledge transfer
Nikolaos Passalis and Anastasios Tefas · 2018
Earlier work this paper cites.
High-Dimensional Probability: An Introduction with Applications in Data Science
Roman Vershynin · 2018
Earlier work this paper cites.
Deep mutual learning
Ying Zhang, Tao Xiang, Timothy M Hospedales, and Huchuan Lu · 2018
Cited alongside, same era.
Rocket launching: A universal and efficient framework for training well-performing light net
Guorui Zhou, Ying Fan, Runpeng Cui, Weijie Bian, Xiaoqiang Zhu, and Kun Gai · 2018
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Cited alongside, same era.
Input similarity from the neural network perspective
Guillaume Charpiat, Nicolas Girard, Loris Felardos, and Yuliya Tarabalka · 2019
Cited alongside, same era.
On the efficacy of knowledge distillation
Jang Hyun Cho and Bharath Hariharan · 2019
Cited alongside, same era.
Searching for mobilenetv3
Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, et al · 2019
Rafael Müller, Simon Kornblith, and Geoffrey E. Hinton · 2020
Later among the works it cites.
Understanding and improving knowledge distillation
Jiaxi Tang, Rakesh Shivanna, Zhe Zhao, Dong Lin, Anima Singh, Ed H Chi, and Sagar Jain · 2020
Later among the works it cites.
Contrastive representation distillation
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2020
Later among the works it cites.
Revisiting knowledge distillation via label smoothing regularization
Li Yuan, Francis EH Tay, Guilin Li, Tao Wang, and Jiashi Feng · 2020
Later among the works it cites.
Implicit regularization via neural feature alignment
Aristide Baratin, Thomas George, César Laurent, R Devon Hjelm, Guillaume Lajoie, Pascal Vincent, and Simon Lacoste-Julien · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Knowledge distillation via route constrained optimization
Xiao Jin, Baoyun Peng, Yichao Wu, Yu Liu, Jiaheng Liu, Ding Liang, Junjie Yan, and Xiaolin Hu · 2019
Cited alongside, same era.
Sgd on neural networks learns functions of increasing complexity
Dimitris Kalimeris, Gal Kaplun, Preetum Nakkiran, Benjamin Edelman, Tristan Yang, Boaz Barak, and Haofeng Zhang · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
Relational knowledge distillation
Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho · 2019
Cited alongside, same era.
Towards understanding knowledge distillation
Mary Phuong and Christoph Lampert · 2019
Cited alongside, same era.
Similarity-preserving knowledge distillation
Frederick Tung and Greg Mori · 2019
Cited alongside, same era.
Knowledge distillation as semiparametric inference
Tri Dao, Govinda M Kamath, Vasilis Syrgkanis, and Lester Mackey · 2021
Later among the works it cites.
A linearized framework and a new benchmark for model selection for fine-tuning
Aditya Deshpande, Alessandro Achille, Avinash Ravichandran, Hao Li, Luca Zancato, Charless Fowlkes, Rahul Bhotika, Stefano Soatto, and Pietro Perona · 2021
Later among the works it cites.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao · 2021
Later among the works it cites.
Feature kernel distillation
Bobby He and Mete Ozay · 2021
Later among the works it cites.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Like Hui and Mikhail Belkin · 2021
Later among the works it cites.
A statistical perspective on distillation
Aditya K Menon, Ankit Singh Rawat, Sashank Reddi, Seungyeon Kim, and Sanjiv Kumar · 2021
Later among the works it cites.
What can linearized neural networks actually say about generalization?
Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, and Pascal Frossard · 2021
Later among the works it cites.
Follow your path: A progressive method for knowledge distillation
Wenxian Shi, Yuxuan Song, Hao Zhou, Bohan Li, and Lei Li · 2021
Later among the works it cites.
Does knowledge distillation really work?
Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A Alemi, and Andrew G Wilson · 2021
Later among the works it cites.
Knowledge distillation: A good teacher is patient and consistent
Lucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva, Rohan Anil, and Alexander Kolesnikov · 2022
Later among the works it cites.
Knowledge distillation with the reused teacher classifier
Defang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang, Yan Feng, and Chun Chen · 2022
Later among the works it cites.
Better supervisory signals by observing learning paths
Yi Ren, Shangmin Guo, and Danica J. Sutherland · 2022
Later among the works it cites.
Pro-KD: Progressive distillation by following the footsteps of the teacher
Mehdi Rezagholizadeh, Aref Jafari, Puneeth S.M. Saladi, Pranav Sharma, Ali Saheb Pasand, and Ali Ghodsi · 2022
Later among the works it cites.
Learning representation from neural fisher kernel with low-rank approximation
Ruixiang Zhang, Shuangfei Zhai, Etai Littwin, and Joshua M. Susskind · 2022
Later among the works it cites.