Fetching the paper…
Reading the bibliography…
Knowledge distillation, which involves extracting the "dark knowledge" from a teacher network to guide the learning of a student network, has emerged as an essential technique for model compression and transfer learning.
A mathematical theory of communication
C. E. Shannon · 1948
Earlier work this paper cites.
Query by committee
H. S. Seung, M. Opper, and H. Sompolinsky · 1992
Earlier work this paper cites.
A sequential algorithm for training text classifiers
David D. Lewis and William A. Gale · 1994
Earlier work this paper cites.
Toward optimal active learning through sampling estimation of error reduction
Nicholas Roy and Andrew McCallum · 2001
Earlier work this paper cites.
Active hidden markov models for information extraction
Tobias Scheffer, Christian Decomain, and Stefan Wrobel · 2001
Earlier work this paper cites.
Model compression
Cristian Buciluundefined, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Multiple-instance active learning
Burr Settles, Mark Craven, and Soumya Ray · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L. Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, S. Shalev-Shwartz, and O. Shamir · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Efficient machine learning for big data: A review
O. Y. Al-Jarrah, P. D. Yoo, S Muhaidat, G. K. Karagiannidis, and K. Taha · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Cited alongside, same era.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Junho Yim, Donggyu Joo, Jihoon Bae, and Junmo Kim · 2017
Cited alongside, same era.
Knowledge transfer via distillation of activation boundaries formed by hidden neurons
Byeongho Heo, Minsik Lee, Sangdoo Yun, and Jin Young Choi · 2019
Later among the works it cites.
Not all samples are created equal: Deep learning with importance sampling
Angelos Katharopoulos and François Fleuret · 2019
Later among the works it cites.
Knowledge distillation via instance relationship graph
Yufan Liu, Jiajiong Cao, Bing Li, Chunfeng Yuan, Weiming Hu, Yangxi Li, and Yunqiang Duan · 2019
Later among the works it cites.
Zero-shot knowledge distillation in deep networks
Gaurav Kumar Nayak, Konda Reddy Mopuri, Vaisakh Shaj, R. Venkatesh Babu, and Anirban Chakraborty · 2019
Later among the works it cites.
Relational knowledge distillation
Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho · 2019
Later among the works it cites.
Mnasnet: Platform-aware neural architecture search for mobile
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Sergey Zagoruyko and Nikos Komodakis · 2017
Cited alongside, same era.
Self-supervised knowledge distillation using singular value decomposition
Seung Hyun Lee, Dae Ha Kim, and Byung Cheol Song · 2018
Cited alongside, same era.
Paraphrasing complex network: Network compression via factor transfer
Jangho Kim, Seonguk Park, and Nojun Kwak · 2018
Cited alongside, same era.
Few-shot learning of neural networks from scratch by pseudo example optimization
Akisato Kimura, Zoubin Ghahramani, Koh Takeuchi, Tomoharu Iwata, and Naonori Ueda · 2018
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz · 2018
Cited alongside, same era.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun · 2018
Cited alongside, same era.
Variational information distillation for knowledge transfer
Sungsoo Ahn, Shell Xu Hu, Andreas Damianou, Neil D. Lawrence, and Zhenwen Dai · 2019
Cited alongside, same era.
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V. Le · 2019
Later among the works it cites.
Similarity-preserving knowledge distillation
Frederick Tung and Greg Mori · 2019
Later among the works it cites.
Supermix: Supervising the mixing data augmentation
Ali Dabouei, Sobhan Soleymani, Fariborz Taherkhani, and Nasser M. Nasrabadi · 2020
Closest in time.
Few sample knowledge distillation for efficient network compression
Tianhong Li, Jianguo Li, Zhuang Liu, and Changshui Zhang · 2020
Closest in time.
Contrastive representation distillation
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2020
Closest in time.
Neural networks are more productive teachers than human raters: Active mixup for data-efficient knowledge distillation from a blackbox model
Dongdong Wang, Yandong Li, Liqiang Wang, and Boqing Gong · 2020
Closest in time.
Knowledge distillation meets self-supervision
Guodong Xu, Ziwei Liu, Xiaoxiao Li, and Chen Change Loy · 2020
Closest in time.
Knowledge distillation via adaptive instance normalization
Jing Yang, Brais Martinez, Adrian Bulat, and Georgios Tzimiropoulos · 2020
Closest in time.