Fetching the paper…
Reading the bibliography…
In this research, we propose an innovative method to boost Knowledge Distillation efficiency without the need for resource-heavy teacher models.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton · 2012
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition, 2014
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition, 2014
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Earlier work this paper cites.
Tiny imagenet visual recognition challenge
Ya Le and Xuan Yang · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Wide residual networks, 2016
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Residual attention network for image classification, 2017
Fei Wang, Mengqing Jiang, Chen Qian, Shuo Yang, Cheng Li, Honggang Zhang, Xiaogang Wang, and Xiaoou Tang · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Cited alongside, same era.
Memory-efficient implementation of densenets, 2017
Geoff Pleiss, Danlu Chen, Gao Huang, Tongcheng Li, Laurens van der Maaten, and Kilian Q. Weinberger · 2017
Cited alongside, same era.
Paying more attention to attention: improving the performance of convolutional neural networks via attention transfer
Nikos Komodakis and Sergey Zagoruyko · 2017
Cited alongside, same era.
Learning deep representations with probabilistic knowledge transfer
Nikolaos Passalis and Anastasios Tefas · 2018
Cited alongside, same era.
Self-supervised knowledge distillation using singular value decomposition
Training data-efficient image transformers & distillation through attention, 2020
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2020
Later among the works it cites.
A case for soft loss functions
Alexandra Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio · 2020
Later among the works it cites.
Revisiting knowledge distillation via label smoothing regularization
Li Yuan, Francis EH Tay, Guilin Li, Tao Wang, and Jiashi Feng · 2020
Later among the works it cites.
Structured knowledge distillation for dense prediction
Yifan Liu, Changyong Shu, Jingdong Wang, and Chunhua Shen · 2020
Later among the works it cites.
Improved knowledge distillation via teacher assistant
Seyed Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh · 2020
Later among the works it cites.
TinyBERT: Distilling BERT for natural language understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Seung Hyun Lee, Dae Ha Kim, and Byung Cheol Song · 2018
Cited alongside, same era.
Incremental learning through deep adaptation
Amir Rosenfeld and John K Tsotsos · 2018
Cited alongside, same era.
Born again neural networks
Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Cited alongside, same era.
Human uncertainty makes classification more robust
Joshua C Peterson, Ruairidh M Battleday, Thomas L Griffiths, and Olga Russakovsky · 2019
Cited alongside, same era.
Relational knowledge distillation
Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Prue: Distilling knowledge from sparse teacher networks, 2022
Shaopu Wang, Xiaojun Chen, Mengzhen Kou, and Jinqiao Shi · 2022
Later among the works it cites.
Eliciting and learning with soft labels from every annotator
Katherine M Collins, Umang Bhatt, and Adrian Weller · 2022
Later among the works it cites.
Decoupled knowledge distillation
Borui Zhao, Quan Cui, Renjie Song, Yiyu Qiu, and Jiajun Liang · 2022
Later among the works it cites.
Knowledge distillation: A good teacher is patient and consistent
Lucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva, Rohan Anil, and Alexander Kolesnikov · 2022
Later among the works it cites.