Fetching the paper…
Reading the bibliography…
Knowledge distillation is an effective approach to learn compact models (students) with the supervision of large and strong models (teachers).
The information bottleneck method
Naftali Tishby, Fernando C. N. Pereira, and William Bialek · 2000
Earlier work this paper cites.
Model compression
Cristian Bucila, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Network in network
Min Lin, Qiang Chen, and Shuicheng Yan · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2015
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Learning from multiple teacher networks
Shan You, Chang Xu, Chao Xu, and Dacheng Tao · 2017
Earlier work this paper cites.
Snapshot ensembles: Train 1, get M for free
Gao Huang, Yixuan Li, Geoff Pleiss, Zhuang Liu, John E. Hopcroft, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Sergey Zagoruyko and Nikos Komodakis · 2017
Earlier work this paper cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Junho Yim, Donggyu Joo, Ji-Hoon Bae, and Junmo Kim · 2017
Earlier work this paper cites.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Cited alongside, same era.
SGDR: stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Born-again neural networks
Tommaso Furlanello, Zachary Chase Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Cited alongside, same era.
Deep mutual learning
Ying Zhang, Tao Xiang, Timothy M. Hospedales, and Huchuan Lu · 2018
Cited alongside, same era.
Emergence of invariance and disentanglement in deep representations
Alessandro Achille and Stefano Soatto · 2018
Cited alongside, same era.
Knowledge distillation by on-the-fly native ensemble
Xu Lan, Xiatian Zhu, and Shaogang Gong · 2018
Infobot: Transfer and exploration via the information bottleneck
Anirudh Goyal, Riashat Islam, Daniel Strouse, Zafarali Ahmed, Hugo Larochelle, Matthew Botvinick, Yoshua Bengio, and Sergey Levine · 2019
Later among the works it cites.
Variational information distillation for knowledge transfer
Sungsoo Ahn, Shell Xu Hu, Andreas Damianou, Neil D Lawrence, and Zhenwen Dai · 2019
Later among the works it cites.
Online knowledge distillation with diverse peers
Defang Chen, Jian-Ping Mei, Can Wang, Yan Feng, and Chun Chen · 2020
Later among the works it cites.
Online knowledge distillation via multi-branch diversity enhancement
Zheng Li, Ying Huang, Defang Chen, Tianren Luo, Ning Cai, and Zhigeng Pan · 2020
Later among the works it cites.
Improved knowledge distillation via teacher assistant
Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine, Akihiro Matsukawa, and Hassan Ghasemzadeh · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the information bottleneck theory of deep learning
Andrew M. Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D. Tracey, and David D. Cox · 2018
Cited alongside, same era.
Training deep neural networks in generations: A more tolerant teacher educates better students
Chenglin Yang, Lingxi Xie, Siyuan Qiao, and Alan L Yuille · 2019
Cited alongside, same era.
On the efficacy of knowledge distillation
Jang Hyun Cho and Bharath Hariharan · 2019
Cited alongside, same era.
Similarity-preserving knowledge distillation
Frederick Tung and Greg Mori · 2019
Cited alongside, same era.
Correlation congruence for knowledge distillation
Baoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou, Yichao Wu, Jiaheng Liu, Zhaoning Zhang, and Yu Liu · 2019
Cited alongside, same era.
Revisit knowledge distillation: a teacher-free framework
Li Yuan, Francis E. H. Tay, Guilin Li, Tao Wang, and Jiashi Feng · 2019
Cited alongside, same era.
Mengya Gao, Yujun Shen, Quanquan Li, and Chen Change Loy · 2020
Later among the works it cites.
Kernelized information bottleneck leads to biologically plausible 3-factor hebbian learning in deep networks
Roman Pogodin and Peter E. Latham · 2020
Later among the works it cites.
Dynamics generalization via information bottleneck in deep reinforcement learning
Xingyu Lu, Kimin Lee, Pieter Abbeel, and Stas Tiomkin · 2020
Later among the works it cites.
Learning student-friendly teacher networks for knowledge distillation
Dae Young Park, Moon-Hyun Cha, Daesin Kim, Bohyung Han, et al · 2021
Later among the works it cites.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao · 2021
Later among the works it cites.
Towards learning spatially discriminative feature representations
Chaofei Wang, Jiayu Xiao, Yizeng Han, Qisen Yang, Shiji Song, and Gao Huang · 2021
Later among the works it cites.
Robust deep reinforcement learning via multi-view information bottleneck
Jiameng Fan and Wenchao Li · 2021
Later among the works it cites.
Revisiting locally supervised learning: an alternative to end-to-end training
Yulin Wang, Zanlin Ni, Shiji Song, Le Yang, and Gao Huang · 2021
Later among the works it cites.
Revisiting locally supervised learning: an alternative to end-to-end training
Yulin Wang, Zanlin Ni, Shiji Song, Le Yang, and Gao Huang · 2021
Later among the works it cites.
Tc3kd: Knowledge distillation via teacher-student cooperative curriculum customization
Chaofei Wang, Ke Yang, Shaowei Zhang, Gao Huang, and Shiji Song · 2022
Closest in time.