Fetching the paper…
Reading the bibliography…
Knowledge distillation is the procedure of transferring "knowledge" from a large model (the teacher) to a more compact one (the student), often being used in the context of model compression.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Model compression
Cristian Buciluundefined, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Geoffrey E. Hinton, Simon Osindero, and Yee Whye Teh · 2006
Earlier work this paper cites.
Scaling learning algorithms towards AI
Yoshua Bengio and Yann LeCun · 2007
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Y. Bengio · 2014
Earlier work this paper cites.
Unsupervised domain adaptation by backpropagation, 2015
Yaroslav Ganin and Victor Lempitsky · 2015
Earlier work this paper cites.
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
SSD: single shot multibox detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott E. Reed, Cheng-Yang Fu, and Alexander C. Berg · 2015
Earlier work this paper cites.
Inceptionism: Going deeper into neural networks, Jun 2015
Alexander Mordvintsev · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Learning using privileged information: Similarity control and knowledge transfer
Vladimir Vapnik and Rauf Izmailov · 2015
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Unifying distillation and privileged information
David Lopez-Paz, Léon Bottou, Bernhard Schölkopf, and Vladimir Naumovich Vapnik · 2016
Earlier work this paper cites.
SGDR: stochastic gradient descent with restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromańska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer T. Chayes, Levent Sagun, and Riccardo Zecchina · 2017
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout
Terrance DeVries and Graham W Taylor · 2017
Earlier work this paper cites.
Improved regularization of convolutional neural networks with cutout
Terrance Devries and Graham W. Taylor · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Gintare Karolina Dziugaite and Daniel M. Roy · 2017
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2017
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, and Tom Goldstein · 2017
Cited alongside, same era.
Be your own teacher: Improve the performance of convolutional neural networks via self distillation
Linfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen, Chenglong Bao, and Kaisheng Ma · 2019
Later among the works it cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2020
Later among the works it cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy
Asit K. Mishra and Debbie Marr · 2017
Cited alongside, same era.
Automatic differentiation in pytorch, 2017
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Learning loss for knowledge distillation with conditional adversarial networks
Zheng Xu, Yen-Chang Hsu, and Jiawei Huang · 2017
Cited alongside, same era.
Darkrank: Accelerating deep metric learning via cross sample similarities transfer
Yuntao Chen, Naiyan Wang, and Zhaoxiang Zhang · 2018
Cited alongside, same era.
Autoaugment: Learning augmentation policies from data
Ekin Dogus Cubuk, Barret Zoph, Dandelion Mané, Vijay Vasudevan, and Quoc V. Le · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Born again neural networks, 2018
Tommaso Furlanello, Zachary C. Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Cited alongside, same era.
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Later among the works it cites.
DE-RRD: A knowledge distillation framework for recommender system
SeongKu Kang, Junyoung Hwang, Wonbin Kweon, and Hwanjo Yu · 2020
Later among the works it cites.
Ensemble distillation for robust model fusion in federated learning
Tao Lin, Lingjing Kong, Sebastian U. Stich, and Martin Jaggi · 2020
Later among the works it cites.
Self-distillation amplifies regularization in hilbert space, 2020
Hossein Mobahi, Mehrdad Farajtabar, and Peter L. Bartlett · 2020
Later among the works it cites.
Entropic gradient descent algorithms and wide flat minima
Fabrizio Pittorino, Carlo Lucibello, Christoph Feinauer, Enrico M. Malatesta, Gabriele Perugini, Carlo Baldassi, Matteo Negri, Elizaveta Demyanenko, and Riccardo Zecchina · 2020
Later among the works it cites.
Contrastive representation distillation, 2020
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2020
Later among the works it cites.
Self-distillation as instance-specific label smoothing
Zhilu Zhang and Mert R. Sabuncu · 2020
Later among the works it cites.
Deep semi-supervised knowledge distillation for overlapping cervical cell instance segmentation
Yanning Zhou, Hao Chen, Huangjing Lin, and Pheng-Ann Heng · 2020
Later among the works it cites.
Even your teacher needs guidance: Ground-truth targets dampen regularization imposed by self-distillation, 2021
Kenneth Borup and Lars N. Andersen · 2021
Later among the works it cites.
Domain generalization needs stochastic weight averaging for robustness on domain shifts
Junbum Cha, Hancheol Cho, Kyungjae Lee, Seunghyun Park, Yunsung Lee, and Sungrae Park · 2021
Later among the works it cites.
Mosaicking to distill: Knowledge distillation from out-of-domain data, 2021
Gongfan Fang, Yifan Bao, Jie Song, Xinchao Wang, Donglin Xie, Chengchao Shen, and Mingli Song · 2021
Later among the works it cites.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen J. Maybank, and Dacheng Tao · 2021
Later among the works it cites.
The many faces of robustness: A critical analysis of out-of-distribution generalization
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer · 2021
Later among the works it cites.
Natural adversarial examples, 2021
Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song · 2021
Later among the works it cites.
Topology distillation for recommender system
SeongKu Kang, Junyoung Hwang, Wonbin Kweon, and Hwanjo Yu · 2021
Later among the works it cites.
Slow learning and fast inference: Efficient graph similarity computation via knowledge distillation
Can Qin, Handong Zhao, Lichen Wang, Huan Wang, Yulun Zhang, and Yun Fu · 2021
Later among the works it cites.
Does knowledge distillation really work?
Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi, and Andrew Gordon Wilson · 2021
Later among the works it cites.