Fetching the paper…
Reading the bibliography…
Despite the empirical success of knowledge distillation, current state-of-the-art methods are computationally expensive to train, which makes them difficult to adopt in practice.
On Measures of Entropy and Information
Alfréd Rényi · 1960
Earlier work this paper cites.
Infinitely Divisible Matrices
Rajendra Bhati · 1969
Earlier work this paper cites.
QR decomposition on GPUs
Andrew Kerr, Dan Campbell, and Mark Richards · 2009
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Nonnegative Decomposition of Multivariate Information
Paul L. Williams and Randall D. Beer · 2010
Earlier work this paper cites.
Face recognition using kernel entropy component analysis
B. H. Shekar, M. Sharmila Kumari, Leonid M. Mestetskiy, and Natalia F. Dyshkant · 2011
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Geoffrey Hinton, Sutskever Ilya, James Martens, and George Dahl · 2013
Earlier work this paper cites.
Information theoretic learning with infinitely divisible kernels
Luis G. Sanchez Giraldo and Jose C. Principe · 2013
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2014
Earlier work this paper cites.
ResNet - Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
FitNets: Hints For Thin Deep Nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2015
Earlier work this paper cites.
Measures of entropy from data using infinitely divisible Kernels
Luis Gonzalo Sanchez Giraldo, Murali Rao, and Jose C. Principe · 2015
Earlier work this paper cites.
Very Deep Convolutional Networks For Large-scale Image Recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi · 2016
Earlier work this paper cites.
Wide Residual Networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
Zehao Huang and Naiyan Wang · 2017
Earlier work this paper cites.
Pruning Filters For Efficient Convnets
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf · 2017
Earlier work this paper cites.
Rényi Differential Privacy
Ilya Mironov · 2017
Cited alongside, same era.
FINN: A framework for fast, scalable binarized neural network inference
Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip Leong, Magnus Jahre, and Kees Vissers · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning
Junho Yim · 2017
Cited alongside, same era.
MobileNetV2: Inverted Residuals and Linear Bottlenecks
Michael H. Fox, Kyungmee Kim, and David Ehrenkrantz · 2018
Cited alongside, same era.
Paraphrasing complex network: Network compression via factor transfer
Jangho Kim, Seong Uk Park, and Nojun Kwak · 2018
Contrastive representation distillation
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2019
Later among the works it cites.
Similarity-preserving knowledge distillation
Fred Tung and Greg Mori · 2019
Later among the works it cites.
Universally Slimmable Networks and Improved Training Techniques
Jiahui Yu and Thomas Huang · 2019
Later among the works it cites.
Understanding autoencoders with information theoretic concepts
Shujian Yu and José C. Príncipe · 2019
Later among the works it cites.
Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer
Sergey Zagoruyko and Nikos Komodakis · 2019
Later among the works it cites.
Wasserstein Contrastive Representation Distillation
Liqun Chen, Dong Wang, Zhe Gan, Jingjing Liu, Ricardo Henao, and Lawrence Carin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning Deep Representations with Probabilistic Knowledge Transfer
Nikolaos Passalis and Anastasios Tefas · 2018
Cited alongside, same era.
MnasNet: Platform-Aware Neural Architecture Search for Mobile
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V. Le · 2018
Cited alongside, same era.
Slimmable Neural Networks
Jiahui Yu, Linjie Yang, Ning Xu, Jianchao Yang, and Thomas Huang · 2018
Cited alongside, same era.
ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
Xiangyu Zhang, Xinyu Zhou, and Mengxiao Lin · 2018
Cited alongside, same era.
Discrimination-aware Channel Pruning for Deep Neural Networks
Zhuangwei Zhuang, Mingkui Tan, Bohan Zhuang, Jing Liu, Yong Guo, Qingyao Wu, Junzhou Huang, and Jinhui Zhu · 2018
Cited alongside, same era.
On the information bottleneck theory of deep learning
Madhu Advani, Artemy Kolchinsky, and Brendan D Tracey · 2019
Cited alongside, same era.
TextBrewer: An Open-Source Knowledge Distillation Toolkit for Natural Language Processing
Yang et al · 2020
Later among the works it cites.
Reducing the Teacher-Student Gap via Spherical Knowledge Distillation
Jia Guo, Minghao Chen, Yao Hu, Chen Zhu, Xiaofei He, and Deng Cai · 2020
Later among the works it cites.
Rotated binary neural network
Mingbao Lin, Rongrong Ji, Zihan Xu, Baochang Zhang, Yan Wang, Yongjian Wu, Feiyue Huang, and Chia Wen Lin · 2020
Later among the works it cites.
Cascaded channel pruning using hierarchical self-distillation
Roy Miles and Krystian Mikolajczyk · 2020
Later among the works it cites.
Forward and Backward Information Retention for Accurate Binary Neural Networks
Haotong Qin, Ruihao Gong, Xianglong Liu, Mingzhu Shen, Ziran Wei, Fengwei Yu, and Jingkuan Song · 2020
Later among the works it cites.
Knowledge Distillation Meets Self-supervision
Guodong Xu, Ziwei Liu, Xiaoxiao Li, and Chen Change Loy · 2020
Later among the works it cites.
Understanding Convolutional Neural Networks With Information Theory: An Initial Exploration
Shujian Yu, Kristoffer Wickstrom, Robert Jenssen, and Jose C. Principe · 2020
Later among the works it cites.
Exploring Inter-Channel Correlation for Diversity-preserved Knowledge Distillation
Li Liu, Qingle Huang, Sihao Lin, Hongwei Xie, Bing Wang, Xiaojun Chang, and Xiaodan Liang · 2021
Closest in time.
ReCU: Reviving the Dead Weights in Binary Neural Networks
Zihan Xu, Mingbao Lin, Jianzhuang Liu, Jie Chen, Ling Shao, Yue Gao, Yonghong Tian, and Rongrong Ji · 2021
Closest in time.
Hierarchical Self-supervised Augmented Knowledge Distillation
Chuanguang Yang, Zhulin An, Linhang Cai, and Yongjun Xu · 2021
Closest in time.
Deep deterministic information bottleneck with matrix-based entropy functional
Xi Yu, Shujian Yu, and José C. Príncipe · 2021
Closest in time.
Barlow Twins: Self-Supervised Learning via Redundancy Reduction
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny · 2021
Closest in time.