Fetching the paper…
Reading the bibliography…
We investigate the mechanisms of self-distillation in multi-class classification, particularly in the context of linear probing with fixed feature extractors where traditional feature learning explanations do not apply.
The rotation of eigenvectors by a perturbation. iii
Chandler Davis and William Morton Kahan · 1970
Earlier work this paper cites.
Caltech-256 object category dataset
Gregory Griffin, Alex Holub, and Pietro Perona · 2007
Earlier work this paper cites.
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Learning from partial labels
Timothee Cour, Ben Sapp, and Ben Taskar · 2011
Earlier work this paper cites.
3d object representations for fine-grained categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei · 2013
Earlier work this paper cites.
Food-101–mining discriminative components with random forests
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Unifying distillation and privileged information
David Lopez-Paz, Léon Bottou, Bernhard Schölkopf, and Vladimir Vapnik · 2015
Earlier work this paper cites.
Distilling knowledge from ensembles of neural networks for speech recognition
Yevgen Chebotar and Austin Waters · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Knowledge distillation across ensembles of multilingual models for low-resource languages
Jia Cui, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Tom Sercu, Kartik Audhkhasi, Abhinav Sethy, Markus Nussbaum-Thom, and Andrew Rosenberg · 2017
Cited alongside, same era.
Born again neural networks
Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Cited alongside, same era.
Generalized cross entropy loss for training deep neural networks with noisy labels
Zhilu Zhang and Mert Sabuncu · 2018
Cited alongside, same era.
Ensemble knowledge distillation for learning improved and efficient networks
Umar Asif, Jianbin Tang, and Stefan Harrer · 2019
Cited alongside, same era.
Partial label learning via label enhancement
Ning Xu, Jiaqi Lv, and Xin Geng · 2019
Later among the works it cites.
Be your own teacher: Improve the performance of convolutional neural networks via self distillation
Linfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen, Chenglong Bao, and Kaisheng Ma · 2019
Later among the works it cites.
Knowledge distillation in wide neural networks: Risk bound, data efficiency and imperfect teacher
Guangda Ji and Zhanxing Zhu · 2020
Later among the works it cites.
Progressive identification of true labels for partial-label learning
Jiaqi Lv, Miao Xu, Lei Feng, Gang Niu, Xin Geng, and Masashi Sugiyama · 2020
Later among the works it cites.
Self-distillation amplifies regularization in hilbert space
Hossein Mobahi, Mehrdad Farajtabar, and Peter Bartlett · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bin Dong, Jikai Hou, Yiping Lu, and Zhihua Zhang · 2019
Cited alongside, same era.
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao · 2019
Cited alongside, same era.
Gm-pll: Graph matching based partial label learning
Gengyu Lyu, Songhe Feng, Tao Wang, Congyan Lang, and Yidong Li · 2019
Cited alongside, same era.
Towards understanding knowledge distillation
Mary Phuong and Christoph Lampert · 2019
Cited alongside, same era.
Contrastive representation distillation
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2019
Cited alongside, same era.
Adaptive graph guided disambiguation for partial label learning
Deng-Bao Wang, Li Li, and Min-Ling Zhang · 2019
Cited alongside, same era.
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Pico: Contrastive label disambiguation for partial label learning
Haobo Wang, Ruixuan Xiao, Yixuan Li, Lei Feng, Gang Niu, Gang Chen, and Junbo Zhao · 2021
Later among the works it cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2022
Later among the works it cites.
Understanding self-distillation in the presence of label noise
Rudrajit Das and Sujay Sanghavi · 2023
Later among the works it cites.