Fetching the paper…
Reading the bibliography…
It has been recently demonstrated that multi-generational self-distillation can improve generalization.
Nonparametric entropy estimation: An overview
Jan Beirlant, Edward J Dudewicz, László Györfi, and Edward C Van der Meulen · 1997
Earlier work this paper cites.
Estimating a dirichlet distribution, 2000
Thomas Minka · 2000
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Caltech-UCSD Birds 200
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona · 2010
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2014
Earlier work this paper cites.
Bayesian dark knowledge
Anoop Korattikara Balan, Vivek Rathod, Kevin P Murphy, and Max Welling · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Unifying distillation and privileged information
David Lopez-Paz, Léon Bottou, Bernhard Schölkopf, and Vladimir Vapnik · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush · 2016
Earlier work this paper cites.
Temporal ensembling for semi-supervised learning
Samuli Laine and Timo Aila · 2016
Earlier work this paper cites.
Distillation as a defense to adversarial perturbations against deep neural networks
Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Cited alongside, same era.
Like what you like: Knowledge distill via neuron selectivity transfer
Zehao Huang and Naiyan Wang · 2017
Cited alongside, same era.
Learning without forgetting
Zhizhong Li and Derek Hoiem · 2017
Cited alongside, same era.
Regularizing neural networks by penalizing confident output distributions
Do we train on test data? purging cifar of near-duplicates
Björn Barz and Joachim Denzler · 2019
Later among the works it cites.
Data-free learning of student networks
Hanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang, Chuanjian Liu, Boxin Shi, Chunjing Xu, Chao Xu, and Qi Tian · 2019
Later among the works it cites.
On the efficacy of knowledge distillation
Jang Hyun Cho and Bharath Hariharan · 2019
Later among the works it cites.
A comprehensive overhaul of feature distillation
Byeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park, Nojun Kwak, and Jin Young Choi · 2019
Later among the works it cites.
Ensemble distribution distillation
Andrey Malinin, Bruno Mlodozeniec, and Mark Gales · 2019
Later among the works it cites.
Zero-shot knowledge transfer via adversarial belief matching
Paul Micaelli and Amos J Storkey · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gabriel Pereyra, George Tucker, Jan Chorowski, Łukasz Kaiser, and Geoffrey Hinton · 2017
Cited alongside, same era.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Antti Tarvainen and Harri Valpola · 2017
Cited alongside, same era.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Junho Yim, Donggyu Joo, Jihoon Bae, and Junmo Kim · 2017
Cited alongside, same era.
Visual relationship detection with internal and external linguistic knowledge distillation
Ruichi Yu, Ang Li, Vlad I Morariu, and Larry S Davis · 2017
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Cited alongside, same era.
Maximum-entropy fine grained classification
Abhimanyu Dubey, Otkrist Gupta, Ramesh Raskar, and Nikhil Naik · 2018
Cited alongside, same era.
Born again neural networks
Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Cited alongside, same era.
Later among the works it cites.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E Hinton · 2019
Later among the works it cites.
Human uncertainty makes classification more robust
Joshua C Peterson, Ruairidh M Battleday, Thomas L Griffiths, and Olga Russakovsky · 2019
Later among the works it cites.
Towards understanding knowledge distillation
Mary Phuong and Christoph Lampert · 2019
Later among the works it cites.
Training deep neural networks in generations: A more tolerant teacher educates better students
Chenglin Yang, Lingxi Xie, Siyuan Qiao, and Alan L Yuille · 2019
Later among the works it cites.
Knowledge extraction with no observable data
Jaemin Yoo, Minyong Cho, Taebum Kim, and U Kang · 2019
Later among the works it cites.
Revisit knowledge distillation: a teacher-free framework
Li Yuan, Francis EH Tay, Guilin Li, Tao Wang, and Jiashi Feng · 2019
Later among the works it cites.
Understanding knowledge distillation in non-autoregressive machine translation
Chunting Zhou, Graham Neubig, and Jiatao Gu · 2019
Later among the works it cites.
Zeroq: A novel zero shot quantization framework
Yaohui Cai, Zhewei Yao, Zhen Dong, Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2020
Closest in time.
Yichen Shen, Zhilu Zhang, Mert R Sabuncu, and Lin Sun · 2020
Closest in time.