Fetching the paper…
Reading the bibliography…
Self-Distillation is a special type of knowledge distillation where the student model has the same architecture as the teacher model.
Do deep nets really need to be deep?
J. Ba and R. Caruana · 2014
Earlier work this paper cites.
Airfoil Self-Noise
T. Brooks, D. Pope, and M. Marcolini · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets; 2014
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
Unifying distillation and privileged information
D. Lopez-Paz, L. Bottou, B. Schölkopf, and V. Vapnik · 2015
Earlier work this paper cites.
Distillation as a defense to adversarial perturbations against deep neural networks
N. Papernot, P. McDaniel, X. Wu, S. Jha, and A. Swami · 2016
Earlier work this paper cites.
Policy distillation
A. A. Rusu, S. G. Colmenarejo, C. Gulcehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell · 2016
Earlier work this paper cites.
Air Quality
S. Vito · 2016
Earlier work this paper cites.
Appliances Energy Prediction
L. Candanedo · 2017
Earlier work this paper cites.
Learning efficient object detection models with knowledge distillation
G. Chen, W. Choi, X. Yu, T. Han, and M. Chandraker · 2017
Earlier work this paper cites.
Like what you like: Knowledge distill via neuron selectivity transfer
Z. Huang and N. Wang · 2017
Earlier work this paper cites.
Learning from noisy labels with distillation
Y. Li, J. Yang, Y. Song, L. Cao, J. Luo, and L.-J. Li · 2017
Earlier work this paper cites.
Do deep convolutional nets really need to be deep and convolutional?
G. Urban, K. J. Geras, S. E. Kahou, O. Aslan, S. Wang, A. Mohamed, M. Philipose, M. Richardson, and R. Caruana · 2017
Earlier work this paper cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
J. Yim, D. Joo, J. Bae, and J. Kim · 2017
Earlier work this paper cites.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
S. Zagoruyko and N. Komodakis · 2017
Cited alongside, same era.
Born again neural networks
T. Furlanello, Z. Lipton, M. Tschannen, L. Itti, and A. Anandkumar · 2018
Cited alongside, same era.
Improving the interpretability of deep neural networks with knowledge distillation
X. Liu, X. Wang, and S. Matwin · 2018
Cited alongside, same era.
Deep mutual learning
Y. Zhang, T. Xiang, T. M. Hospedales, and H. Lu · 2018
Cited alongside, same era.
Variational information distillation for knowledge transfer
S. Ahn, S. X. Hu, A. Damianou, N. D. Lawrence, and Z. Dai · 2019
Cited alongside, same era.
Structured knowledge distillation for semantic segmentation
Y. Liu, K. Chen, C. Liu, Z. Qin, Z. Luo, and J. Wang · 2019
Cited alongside, same era.
Emerging properties in self-supervised vision transformers
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin · 2021
Later among the works it cites.
Knowledge distillation: A survey
J. Gou, B. Yu, S. J. Maybank, and D. Tao · 2021
Later among the works it cites.
Neural attention distillation: Erasing backdoor triggers from deep neural networks, 2021
Y. Li, X. Lyu, N. Koren, L. Lyu, B. Li, and X. Ma · 2021
Later among the works it cites.
A statistical perspective on distillation
A. K. Menon, A. S. Rawat, S. Reddi, S. Kim, and S. Kumar · 2021
Later among the works it cites.
Self-distillation: Towards efficient and compact neural networks
L. Zhang, C. Bao, and K. Ma · 2021
Later among the works it cites.
Teacher-student architecture for knowledge learning: A survey
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards understanding knowledge distillation
M. Phuong and C. Lampert · 2019
Cited alongside, same era.
Patient knowledge distillation for bert model compression
S. Sun, Y. Cheng, Z. Gan, and J. Liu · 2019
Cited alongside, same era.
Be your own teacher: Improve the performance of convolutional neural networks via self distillation
L. Zhang, J. Song, A. Gao, J. Chen, C. Bao, and K. Ma · 2019
Cited alongside, same era.
The optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization
D. Kobak, J. Lomond, and B. Sanchez · 2020
Cited alongside, same era.
Self-distillation amplifies regularization in hilbert space
H. Mobahi, M. Farajtabar, and P. Bartlett · 2020
Cited alongside, same era.
Countermeasure against backdoor attack on neural networks utilizing knowledge distillation
K. Yoshida and T. Fujino · 2020
Cited alongside, same era.
C. Hu, X. Li, D. Liu, X. Chen, J. Wang, and X. Liu · 2022
Later among the works it cites.
Eliminating backdoor triggers for deep neural networks using attention relation graph distillation, 2022
J. Xia, T. Wang, J. Ding, X. Wei, and M. Chen · 2022
Later among the works it cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Z. Allen-Zhu and Y. Li · 2023
Later among the works it cites.
Understanding self-distillation in the presence of label noise
R. Das and S. Sanghavi · 2023
Later among the works it cites.
Label poisoning is all you need
R. Jha, J. Hayase, and S. Oh · 2023
Later among the works it cites.
H. Jeong and H. W. Chung · 2024
Closest in time.
Rho-1: Not all tokens are what you need
Z. Lin, Z. Gou, Y. Gong, X. Liu, Y. Shen, R. Xu, C. Lin, Y. Yang, J. Jiao, N. Duan, et al · 2024
Closest in time.
Iterative data smoothing: Mitigating reward overfitting and overoptimization in rlhf
B. Zhu, M. I. Jordan, and J. Jiao · 2024
Closest in time.